Reading additional input from stdin... OpenAI Codex v0.149.0 -------- workdir: /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl model: gpt-5.6-sol provider: openai approval: never sandbox: workspace-write [workdir, /tmp, $TMPDIR, /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm, /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.git, /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.git, /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001] reasoning effort: high reasoning summaries: none session id: 01a032d1-af70-7903-99d2-14de6b7620fe -------- user You are the sole Codex worker for Backgammon Simplified research lane explainer-native-cubeful-k001. Read /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/milestones/explainer-native-cubeful-k001/handoffs/current.md and every referenced frozen protocol before acting. Execute exactly the current frozen experiment autonomously through durable artifacts, tests, result, and Git commits. Adapt only between completed experiments, never within a running protocol. Existing accepted data only. No new GNU, Sage, matches, labels, generic 0-ply, Sage/GNU campaign training consumption, production promotion, Analyzer changes, Canonical changes, or Corpus changes. Preserve TRAIN, DEVELOPMENT/ADAPTIVE SELECTION, PROTECTED FINAL EVALUATION and record every protected access. Accepted architecture remains ridge-ranking-hadd-value-explanation-sidecar-v1. Any better model is only CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION. On HFCS, Corpus and Post Match have priority and substantial headroom must remain. Work only in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl, /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm, /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime and /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/artifacts. Commit results locally but do not push; the supervisor pushes after verifying clean state. At completion update the lane handoff to completed/waiting-for-Research-Director and do not invent the next experiment yourself. codex I’ll first read the lane handoff, repository guidance, and every frozen protocol it references. Then I’ll execute only that protocol, preserve the evaluation partitions and access log, verify the durable artifacts, commit locally, and leave the lane waiting for the Research Director. exec /bin/bash -lc "pwd && rg --files -g 'AGENTS.md' -g 'current.md' -g '*protocol*' -g '*.md' .. | sed -n '1,240p'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 117ms: /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl ../impl/TASK_RESEARCH.md ../impl/docs/deep-4ply-acquisition-v3.md ../impl/docs/modeling/feature-v2-100-experiment-v1.md ../impl/docs/modeling/evaluation-harness-v2.md ../impl/docs/modeling/explainer-k002-hadd-compact-runtime-v1.md ../impl/docs/modeling/feature-v2-250-experiment-v1.md ../impl/docs/modeling/feature-v2-sidecars-v1.md ../impl/docs/modeling/improve-additive-native-cubeful-modeling-v1.md ../impl/docs/modeling/feature-v2-alternate-model-comparison-v1.md ../impl/docs/modeling/feature-v2-model-selection-v1.md ../impl/docs/modeling/feature-v2-targeted-feature-interaction-v1.md ../impl/docs/modeling/diagnose-native-cubeful-context-error-structure-v1.md ../impl/docs/modeling/retained-actual-4ply-reconciliation-v1.md ../impl/docs/modeling/feature-v2-deep-label-data-scaling-plan-v1.md ../impl/docs/modeling/feature-v2-deep-4ply-modeling-v1.md ../impl/docs/modeling/feature-v2-capacity-test-v1.md ../impl/docs/modeling/position-value-modeling-v1.md ../impl/docs/modeling/feature-v2-502-experiment-v1.md ../impl/docs/modeling/feature-v2-alternate-model-comparison-plan-v1.md ../impl/docs/modeling/canonical-parquet-query-patterns-v1.md ../impl/docs/modeling/feature-v2-shallow-to-deep-protocol-v1.md ../impl/docs/modeling/feature-v2-deep-label-data-efficiency-v1.md ../impl/docs/modeling/explainer-k002-constrained-additive-position-model-v1.md ../impl/docs/modeling/feature-v2-shallow-to-deep-v1.md ../impl/docs/match-context-diagnostic-v2.md ../impl/docs/deep-4ply-acquisition-freeze-v1.md ../impl/docs/contracts/explainer-k002-compact-hadd-integration-contract-v1.md ../impl/docs/contracts/explainer-model-output-parquet-v1.md ../impl/docs/contracts/canonical-analysis-parquet-v1.md ../impl/docs/contracts/match-context-feasibility-v1.md ../impl/docs/contracts/gnu-review-candidate-pilot-schema.md ../impl/docs/contracts/explainer-k002-hadd-analyzer-read-only-integration-v1.md ../impl/docs/contracts/canonical-analysis-parquet-v1-consumer-guide.md ../impl/docs/deep-4ply-expanded-authority-v2.md ../impl/docs/deep-4ply-acquisition-planner-reconciliation-v2.md ../impl/docs/deep-4ply-hfcs-capacity-v1.md ../tm/README.md ../impl/docs/handoffs/research/2026-07-22-match-context-diagnostic-v2.md ../impl/docs/handoffs/research/2026-07-22-candidate-parser-pilot.md ../impl/docs/handoffs/research/LATEST.md ../impl/docs/handoffs/research/2026-07-22-match-context-feasibility.md ../tm/docs/autopilot-v1.md ../tm/scripts/operator/milestone-implementer-bootstrap.md ../tm/tasks/validate-published-engine-kit-wheel-clean-env/README.md ../tm/coordination/control-tower-minion-kickoff-20260816.md ../tm/coordination/current-project-state.md ../tm/coordination/control-tower-replan-20260816.md ../tm/coordination/task-manager-current.md ../tm/milestones/gnuraw-k001/README.md ../tm/tasks/recover-gnu-full-corpus-canonical-conversion-v1/README.md ../tm/milestones/node-k001/assignment.md ../tm/milestones/node-k001/README.md ../tm/milestones/explainer-native-cubeful-k001/prompts/001-improve-additive-native-cubeful-modeling-v1.md ../tm/milestones/node-k001/handoffs/current.md ../tm/milestones/explainer-native-cubeful-k001/prompts/002-diagnose-native-cubeful-context-error-structure-v1.md ../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md ../tm/milestones/explainer-native-cubeful-k001/handoffs/current.md ../tm/milestones/analyzer-k001/README.md ../tm/milestones/retcorpus-k001/prompts/001-lifecycle-v5-kickoff.md ../tm/milestones/retcorpus-k001/prompts/002-hfcs-canonical-writer-baseline.md ../tm/milestones/retcorpus-k001/README.md ../tm/milestones/retcorpus-k001/gateway/shared-server-gateway.md ../tm/milestones/explainer-k002/assignment.md ../tm/milestones/explainer-k002/handoffs/current.md ../tm/milestones/retcorpus-k001/handoffs/current.md ../tm/milestones/retcorpus-k001/assignment.md ../impl/TASK.md ../runtime/python311/usr/lib/python3.11/asyncio/protocols.py ../runtime/python311/usr/lib/python3.11/asyncio/__pycache__/protocols.cpython-311.pyc ../impl/results/REPORT.md ../impl/artifacts/derived/sage_gnu_match_context_diagnostic_v2/README.md ../impl/artifacts/development/explainer-k002-position-value-modeling/protocol.md ../impl/artifacts/development/candidate_parser_pilot/README.md ../impl/artifacts/derived/sage_gnu_match_context_feasibility_v1/README.md ../impl/artifacts/development/explainer-k002-constrained-additive-position-model/protocol-tests.json ../impl/artifacts/development/explainer-k002-constrained-additive-position-model/protocol.md ../impl/artifacts/development/explainer-k002-hadd-compact-runtime/protocol-tests.json ../impl/artifacts/development/explainer-k002-compact-hadd-integration-contract-v1/phase1-selection-review.md ../impl/artifacts/development/explainer-k002-compact-hadd-integration-contract-v1/analyzer-read-only-integration-contract.md ../impl/artifacts/development/explainer-k002-compact-hadd-integration-contract-v1/deployment-contract.md ../impl/artifacts/development/explainer-k002-shallow-to-deep/explainer-k002-shallow-to-deep-v1-9deed0dafe29b2ff/protocol.md exec /bin/bash -lc "rg --files -uu .. | rg '("'^|/)AGENTS'"\\.md"'$|tm/milestones/explainer-native-cubeful-k001/handoffs/current'"\\.md"'$'"' && sed -n '1,260p' ../tm/milestones/explainer-native-cubeful-k001/handoffs/current.md" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 104ms: ../tm/milestones/explainer-native-cubeful-k001/handoffs/current.md # Explainer Native Cubeful K001 Current Handoff Version: `project-task-manager-explainer-native-cubeful-k001-handoff-v5` Updated: 2026-08-24 EDT Status: `GENERATION_3_FROZEN_READY_FOR_IMMEDIATE_CODEX_AFTER_LIVE_PREFLIGHT` Current task: `test-crawford-state-context-family-ablation-v1` Normative frozen protocol: `milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md` Protocol freeze commit: `222c2a50418082c65fd57792564c1dc1cc1ad0f6` Implementation branch: `backgammonsimplified/backgammon-explainer@research/explainer-native-cubeful-k001` Generation 3 starting implementation: `75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a` Accepted product architecture remains fixed: `ridge-ranking-hadd-value-explanation-sidecar-v1`. Pairwise Ridge remains sole recommendation/ranking authority. Compact HADD remains accepted probability/value/explanation authority. Any improved model is only `CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION`. ## Completed Generation 2 evidence Generation 2 `diagnose-native-cubeful-context-error-structure-v1` completed at implementation: `75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a`. Result identity: `b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb`. Artifact package identity: `d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7`. Terminal classification: `CONTEXT_FAILURE_LOCALIZED`. Deterministic routing cell: `crawford_state/money_or_zero`. The primary Generation 1 P3-plus-context stream had eight supported `STABLE_CONTEXT_REGRESSION` cells across all six frozen families. The strongest useful evidence for routing is that the full context candidate's DEVELOPMENT regression remains broad, while the prospectively frozen deterministic localization rule selected the Crawford-state `money_or_zero` cell. Important limitation: every DEVELOPMENT row is money play. Match-length, score-relation, and Crawford-state therefore each have exactly one populated supported `money_or_zero` cell, each with within-family positive squared-error regression mass share `1.0` and identical RMSE regression `+0.027784317582417117`. The selected Crawford family is therefore a deterministic routing result, not causal evidence by itself. Generation 2 performed model fits `0`, used modeling candidates `NONE`, accessed PROTECTED FINAL EVALUATION exactly `0` times, reproduced the exact Generation 1 population and fixed prediction streams, and passed package/hash verification. Production, Analyzer, Canonical, Corpus, Ridge/HADD roles, and calculated-cubeful authority remained unchanged. ## Frozen Generation 3 question The Generation 2 protocol prospectively permits one bounded context-family representation or ablation tied to the selected routing family. Generation 3 therefore performs exactly one causal test: - reproduce the Generation 1 full P3-plus-context additive specification; - remove only accepted context field `cubeful_crawford` before additive basis construction; - retain all other 14 context fields and all P3 position features unchanged; - reuse exact TRAIN/DEVELOPMENT memberships, four grouped folds, target semantics, additive basis, alpha `100.0`, optimizer, seed, and numerical policy from Generation 1; - fit exactly one new candidate; - perform no hyperparameter search, feature search, interaction search, or alternative-family ablation; - PROTECTED FINAL EVALUATION access budget remains exactly `0`. The terminal rule is prospectively frozen. `CRAWFORD_CONTEXT_FAMILY_CAUSAL` requires at least `0.005000` global RMSE recovery and `0.002000` global MAE recovery versus the Generation 1 full-context control, stable improvement in at least `3/4` grouped folds, no fold regression above `0.002000`, and no material global-bias regression. Otherwise the result is `CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL`. No result in this generation changes production or automatically selects a model. ## Coordinator execution The existing HFCS lane supervisor should consume this handoff and launch one Codex worker after a fresh host preflight. Exactly one active Codex is permitted for this lane. Before substantive work, verify fresh CPU/RAM/disk/process safety. Run the Explainer task at nice `19` with conservative numerical threading. Corpus Listener/writer, Post Match, and existing GNU/Sage processes retain priority. Never signal, stop, renice, restart, reconfigure, or otherwise mutate those priority processes. Execute only the frozen Generation 3 protocol through durable implementation artifacts, tests, result, and Task Management closeout. At completion leave the lane `COMPLETED_WAITING_FOR_RESEARCH_DIRECTOR`; do not invent Generation 4 within the Codex worker. ## Absolute boundaries Existing accepted data only. New GNU `0`; new Sage `0`; new matches `0`; new labels `0`; generic new 0-ply `0`; Sage/GNU campaign validation data as training `NO`; production promotion `NO`; Analyzer mutation `NO`; Canonical mutation `NO`; Corpus mutation `NO`; calculated-cubeful authority change `NO`. `CUBEFUL_CALCULATION_AUTHORITY_BLOCKED` remains absolute. exec /bin/bash -lc "wc -l ../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md && sed -n '1,320p' ../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 33ms: 188 ../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md # Frozen Protocol: Test Crawford-State Context Family Ablation v1 Status: `FROZEN_READY_FOR_CODEX` Lane: `explainer-native-cubeful-k001` Task: `test-crawford-state-context-family-ablation-v1` Host: `HFCS` ## Starting authority Implementation branch: `research/explainer-native-cubeful-k001` Starting implementation: `75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a` Accepted product architecture remains fixed: `ridge-ranking-hadd-value-explanation-sidecar-v1`. Pairwise Ridge remains sole recommendation/ranking authority. Compact HADD remains accepted probability/value/explanation authority. Accepted integration package identity: `f40ba9417896383a94e48012843eb0f45e177430e746cd01080f8243c5751424`. `CUBEFUL_CALCULATION_AUTHORITY_BLOCKED` remains absolute. Generation 1 result identity: `6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be`. Generation 1 artifact package identity: `65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7`. Generation 2 result identity: `b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb`. Generation 2 artifact package identity: `d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7`. ## Completed evidence motivating this experiment Generation 1 established that the frozen additive native-Cubeful candidates were worse than the accepted baseline on DEVELOPMENT. Accepted baseline RMSE/MAE were `0.29707332956861665` / `0.2215157712310875`; P3-only additive were `0.312778149409631` / `0.22984733691163486`; P3-plus-context additive were `0.32485764715103377` / `0.23161408367358322`. Neither additive candidate improved both primary metrics in any of the four grouped folds. PROTECTED FINAL EVALUATION was not accessed. Generation 2 prospectively diagnosed the error structure of the fixed Generation 1 streams and returned `CONTEXT_FAILURE_LOCALIZED`. The deterministic routing cell was `crawford_state/money_or_zero`. That routing result has an important frozen limitation: all DEVELOPMENT rows are money play. Therefore match-length, score-relation, and Crawford-state each contain only one populated supported `money_or_zero` cell, each with within-family positive squared-error regression mass share `1.0` and identical RMSE regression `+0.027784317582417117`. The Generation 2 deterministic tie-break selected `crawford_state/money_or_zero`; it did not establish that Crawford state is causally responsible. The Generation 2 frozen follow-up authority permits exactly one bounded context-family representation or ablation tied to the single routing family. This protocol performs the smallest causal test: remove only the Crawford-state context family from the otherwise unchanged Generation 1 P3-plus-context additive model. ## Scientific question Does the accepted Crawford-state context feature family materially contribute to the Generation 1 additive native-Cubeful DEVELOPMENT regression, or was the Generation 2 localization only a consequence of the all-money support geometry and deterministic tie-break? ## Frozen data and target Use exactly the existing Generation 1 partitions and no other rows. TRAIN: - candidates: `20,981,224`; - decisions: `1,000,002`; - complete-game groups: `32,228`; - checkpoint: `1000000`; - membership identity: `2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec`. DEVELOPMENT: - candidates: `2,094,039`; - decisions: `100,015`; - complete-game groups: `3,178`; - membership identity: `5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544`; - grouped folds: exactly the inherited four complete-game folds from Generation 1, assignment `SHA256([version,seed,complete_game_id]) modulo four`, seed `20260823`, version `native-cubeful-development-complete-game-fold-4-v1`. Target authority is unchanged: - id: `native_cubeful_equity_static_next_player`; - evaluation mode: existing GNU native `Cubeful` only; - source field: `native_equity`; - modeled perspective: normalized static post-move next player on roll; - transform: `-native_equity`. Before fitting, reproduce and hash the exact TRAIN and DEVELOPMENT memberships, target mapping, Generation 1 frozen authorities, and Generation 2 result identity. Any mismatch stops with `BLOCKED_INPUT_IDENTITY_MISMATCH`. PROTECTED FINAL EVALUATION access budget: exactly `0`. Protected membership and predictions must not be opened, loaded, or scored. ## Frozen model family Accepted baseline remains descriptive comparison authority: - id: `accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1`; - accepted baseline model identity: `e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5`. Generation 1 full-context additive control: - id: `native-cubeful-p3-context-additive-ridge-v1`; - model identity: `ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69`; - features: P3 position registry plus the complete accepted 15-field factual context block; - position registry: `explainer-position-value-p3-v1-30ede35745bbbc64`; - position registry identity: `30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287`; - context order identity: `ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914`. Exactly one new fitted candidate is permitted: `native-cubeful-p3-context-minus-crawford-additive-ridge-v1` It must be identical to the Generation 1 P3-plus-context additive specification except that the single accepted context field `cubeful_crawford` is removed before basis construction. Remove every additive basis term derived from that field and no other field. Retain all other 14 context fields in their frozen Generation 1 order. No substitute Crawford encoding, no interaction, no new money-play feature, no other field ablation, and no feature search is permitted. ## Frozen basis, regularization, optimizer, and seed Reuse the Generation 1 additive specification exactly: - per-feature terms: standardized linear, hinge q25, hinge q50, hinge q75; - hinge quantiles: `[0.25, 0.50, 0.75]`, derived from TRAIN only; - interactions: `0`; - intercept: `true`; - regularization grid: exactly `[100.0]`; - alpha: exactly `100.0`; - optimizer: deterministic streaming Adam on the Ridge objective; - batch size: `32768`; - beta1: `0.9`; - beta2: `0.999`; - epochs: `12`; - initial learning rate: `0.015`; - learning-rate epoch multiplier: `0.75`; - seed: `20260823`; - numerical threading: one BLAS thread; - training arithmetic: float32 basis/optimizer, float64 retained inference/reconstruction. Hyperparameter search: `0`. Feature search: `0`. Candidate count beyond fixed controls: exactly `1`. Commit a frozen-authorities/input receipt containing the exact resulting 14-field context order and its SHA256 identity before DEVELOPMENT scoring. ## Frozen DEVELOPMENT metrics Report for accepted baseline, Generation 1 P3-plus-context control, and the single Crawford-ablation candidate: - RMSE; - MAE; - signed bias; - correlation/R2 where numerically defined; - exact same metrics in each of the four inherited grouped folds. For the ablation candidate versus Generation 1 full-context control, additionally report: - RMSE recovery = `full_context_rmse - ablation_rmse`; - MAE recovery = `full_context_mae - ablation_mae`; - absolute-bias change; - fold counts with lower RMSE and lower MAE; - maximum fold RMSE regression and maximum fold MAE regression. For the ablation candidate versus accepted baseline, report the same primary metric differences descriptively. The accepted baseline is not refit. ## Prospectively frozen decision rule Classify `CRAWFORD_CONTEXT_FAMILY_CAUSAL` only if the single ablation candidate satisfies all of: 1. global RMSE recovery versus Generation 1 full-context control is at least `0.005000`; 2. global MAE recovery versus Generation 1 full-context control is at least `0.002000`; 3. lower RMSE and lower MAE than the full-context control in at least `3/4` inherited folds; 4. neither RMSE nor MAE regresses versus the full-context control by more than `0.002000` in any fold; 5. absolute global bias does not worsen versus the full-context control by more than `0.005000`. Otherwise classify `CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL`. A secondary descriptive label `ABLATED_MODEL_BEATS_ACCEPTED_BASELINE` may be recorded only if the ablation candidate improves over the accepted baseline by at least `0.002000` in both global RMSE and global MAE, does not worsen absolute global bias by more than `0.005000`, improves both metrics in at least `3/4` folds, and regresses neither metric by more than `0.002000` in any fold. No DEVELOPMENT outcome in this experiment authorizes PROTECTED access. Even if the secondary label is true, the model is only `CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION` and requires a separately frozen future selection protocol before any protected evaluation. ## Frozen routing after completion If `CRAWFORD_CONTEXT_FAMILY_CAUSAL`, the Research Director may next freeze one bounded confirmation or representation simplification specifically tied to the Crawford/money-play context family. Do not search other context families from this experiment. If `CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL`, close the selected Crawford local-patch hypothesis. The Generation 2 `money_or_zero` localization must then be recorded as non-causal under its single-family follow-up test. A later experiment may test a broader native-Cubeful target/model-capacity hypothesis, but must be separately frozen and must not use this failed ablation as permission to sweep other local context families. ## Required durable artifacts Record and hash: - frozen source/partition/target receipt; - exact 14-field context-order receipt; - model/basis/optimizer authority receipt; - candidate model identity and coefficients; - DEVELOPMENT global metrics; - four-fold metrics; - recovery and accepted-baseline comparison tables; - deterministic terminal classification; - protected-access log showing exactly `0`; - validation/test evidence; - package manifest and SHA256SUMS. At completion update the durable result and current handoff, commit implementation and Task Management changes locally, and leave the lane `COMPLETED_WAITING_FOR_RESEARCH_DIRECTOR`. The supervisor owns pushing clean branches. ## HFCS safety Before substantive computation perform a fresh CPU, RAM, disk, and process preflight. Run Explainer work at nice `19` with one numerical worker/thread. Corpus Listener/writer, Post Match, and existing GNU/Sage processes retain priority and must not be signaled, reniced, stopped, restarted, reconfigured, or otherwise mutated. If substantial headroom is not available, stop this lane with a durable `BLOCKED_HOST_SAFETY` result rather than interfering with priority work. ## Absolute boundaries Existing accepted data only. New GNU `0`; new Sage `0`; new matches `0`; new labels `0`; generic new 0-ply `0`; Sage/GNU campaign validation data as training `NO`; production promotion `NO`; Analyzer mutation `NO`; Canonical mutation `NO`; Corpus mutation `NO`; calculated-cubeful authority change `NO`. Accepted Ridge and HADD product roles remain unchanged. Any materially improved research model is only `CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION`. exec /bin/bash -lc "wc -l ../tm/milestones/explainer-native-cubeful-k001/prompts/001-improve-additive-native-cubeful-modeling-v1.md ../tm/milestones/explainer-native-cubeful-k001/prompts/002-diagnose-native-cubeful-context-error-structure-v1.md && sed -n '1,360p' ../tm/milestones/explainer-native-cubeful-k001/prompts/001-improve-additive-native-cubeful-modeling-v1.md && sed -n '1,360p' ../tm/milestones/explainer-native-cubeful-k001/prompts/002-diagnose-native-cubeful-context-error-structure-v1.md" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 42ms: 55 ../tm/milestones/explainer-native-cubeful-k001/prompts/001-improve-additive-native-cubeful-modeling-v1.md 213 ../tm/milestones/explainer-native-cubeful-k001/prompts/002-diagnose-native-cubeful-context-error-structure-v1.md 268 total # Frozen Protocol: Improve Additive Native Cubeful Modeling v1 Status: `FROZEN_READY_FOR_CODEX` Lane: `explainer-native-cubeful-k001` Task: `improve-additive-native-cubeful-modeling-v1` ## Starting authority Implementation branch: `research/explainer-native-cubeful-k001` Starting implementation: `58522bb078ecda273a11476c60f1875a2255b285` Accepted product architecture remains fixed: `ridge-ranking-hadd-value-explanation-sidecar-v1`. Accepted integration package identity: `f40ba9417896383a94e48012843eb0f45e177430e746cd01080f8243c5751424`. `CUBEFUL_CALCULATION_AUTHORITY_BLOCKED` remains absolute. ## Hypothesis The successful nonlinear additive architecture can materially improve prediction of EXISTING GNU native Cubeful equity when given accepted factual position representation plus accepted score/match/cube context, without inventing calculated cubeful logic. ## Data and partition freeze Use existing accepted native Cubeful targets only. Before fitting or inspecting new outcome metrics, discover and record the exact target/data identities and reuse an already frozen TRAIN/DEVELOPMENT/PROTECTED grouping if one exists. If none exists, deterministically create a complete-source-group split from existing accepted data, commit the membership identity, and use that same split for the entire 24-hour program. No Sage/GNU campaign data may be consumed as training. ## Frozen model comparison 1. reproduce the strongest existing accepted native-Cubeful baseline available in durable Explainer evidence; 2. additive model using accepted position representation only; 3. additive model using accepted position representation plus the full accepted factual match/cube context block. Model structure, regularization grid, context fields, preprocessing and seeds must be committed before DEVELOPMENT scoring. No broad search. ## Primary metrics Development RMSE, MAE, bias, correlation/R2 where meaningful, calibration by target magnitude, and exact contribution reconstruction for additive candidates. Segment descriptively by cube ownership, cube value, score, match length, Crawford state and position class using predeclared factual bins. ## Decision rule `MATERIAL_NATIVE_CUBEFUL_IMPROVEMENT` requires a clear development improvement over the reproduced baseline in both RMSE and MAE, no material global-bias regression, and stable direction across grouped folds. A winner is frozen before any PROTECTED access. A protected winner may only be classified `CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION`. Production remains unchanged. If the additive candidate wins, the next separately frozen experiment is context-family ablation and robustness by match state with exact contribution reconstruction. If it remains weak, the next experiment diagnoses cube ownership/value/score/match-length/Crawford/position-class error structure before proposing new context representation. If progress requires missing calculated-cubeful authority, stop that line, record the exact missing interface and redirect the lane. Do not invent it. ## Host boundary Initial host HFCS. Before every substantial run perform fresh CPU/RAM/disk/process preflight. Heavy work runs niced. Preserve substantial headroom. Never signal, stop, renice, restart or reconfigure Post Match, Corpus Listener/writer lease, or historical GNU workloads. Yield only this lane's own process tree. ## Absolute boundaries No new GNU, Sage, source matches, labels, generic 0-ply generation or Sage/GNU campaign training consumption. No production promotion, Analyzer mutation, Canonical mutation or Corpus mutation. Adapt between experiments only. # Frozen Protocol: Diagnose Native Cubeful Context Error Structure v1 Status: `FROZEN_READY_FOR_CODEX` Lane: `explainer-native-cubeful-k001` Task: `diagnose-native-cubeful-context-error-structure-v1` Host: `HFCS` ## Starting authority Implementation branch: `research/explainer-native-cubeful-k001` Starting implementation: `2b2a82284649485b00621cd243dfa17c6accca9c` Accepted product architecture remains fixed: `ridge-ranking-hadd-value-explanation-sidecar-v1`. Accepted integration package identity: `f40ba9417896383a94e48012843eb0f45e177430e746cd01080f8243c5751424`. `CUBEFUL_CALCULATION_AUTHORITY_BLOCKED` remains absolute. Generation 1 result identity: `6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be`. Generation 1 artifact package identity: `65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7`. Generation 1 established that neither frozen additive candidate improved the accepted native-Cubeful baseline on DEVELOPMENT. P3-only and P3-plus-context candidates were worse in both RMSE and MAE and improved both metrics in `0/4` grouped folds. No winner was selected and PROTECTED FINAL EVALUATION was not accessed. The Generation 1 protocol prospectively routes a weak result to a context/error-structure diagnosis before any new context representation is proposed. This protocol executes that route. ## Scientific question Is the Generation 1 additive native-Cubeful failure concentrated in a stable factual cube/match/score/position segment, or is the regression distributed broadly enough that another local context-family patch is not justified? ## Frozen data and predictions Use only the exact existing Generation 1 DEVELOPMENT membership and the exact Generation 1 prediction streams for: 1. `accepted_baseline`; 2. `native-cubeful-p3-additive-ridge-v1`; 3. `native-cubeful-p3-context-additive-ridge-v1`. Exact DEVELOPMENT membership identity: `5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544`. DEVELOPMENT population: `2,094,039` candidates, `100,015` decisions, `3,178` complete-game groups. Reuse exactly the four complete-game grouped DEVELOPMENT folds from Generation 1. Generation 1 TRAIN membership identity remains `2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec`, but TRAIN outcomes are not used in this diagnostic. PROTECTED FINAL EVALUATION access budget: exactly `0`. Protected membership and predictions must not be opened or scored. Before scoring, write an immutable receipt containing all source identities, prediction identities, partition identities, and field mappings. If the frozen DEVELOPMENT membership or Generation 1 prediction identities cannot be reproduced exactly, stop with `BLOCKED_INPUT_IDENTITY_MISMATCH`. ## Model fitting and search Model fitting: `0`. Modeling candidates: `NONE`. Hyperparameter search: `0`. Feature search: `0`. Context feature additions: `0`. This is a diagnostic only. ## Prospectively frozen segment families Compute each family independently. Do not search arbitrary intersections or create outcome-dependent bins. ### 1. Cube ownership Normalize from the existing factual ownership field to exactly: - `centered`; - `on_roll_player_owned`; - `opponent_owned`; - `unavailable_or_unknown`. The mapping from raw accepted values to these four semantic categories must be committed in the input receipt before scoring. ### 2. Cube value Use exactly these bins: - `cube_1`; - `cube_2`; - `cube_4`; - `cube_8`; - `cube_16`; - `cube_32_plus`; - `unavailable_or_unknown`. ### 3. Match length Use exactly: - `money_or_zero` for non-match or length `<= 0`; - `length_1_3`; - `length_4_7`; - `length_8_15`; - `length_16_plus`; - `unavailable_or_unknown`. ### 4. Score relation from modeled next-player perspective For match positions, compute points-away from the existing accepted match length and scores, then classify exactly: - `tied_away`; - `on_roll_ahead` when on-roll player has fewer points away; - `on_roll_behind` when on-roll player has more points away; - `money_or_zero`; - `unavailable_or_unknown`. No finer score bins are permitted in this experiment. ### 5. Crawford state Normalize exactly to: - `crawford`; - `post_crawford` if that state is explicitly available in accepted facts; - `ordinary_match`; - `money_or_zero`; - `unavailable_or_unknown`. Do not infer post-Crawford from outcomes or external data. If no accepted explicit post-Crawford field exists, map those rows to `ordinary_match` and record the field limitation before scoring. ### 6. Position class Reuse the already accepted/frozen factual position-class definitions available to the Explainer research program. Collapse only to exactly: - `race`; - `contact`; - `bearoff`; - `other_or_unknown`. No new classifier may be learned. If an accepted deterministic mapping is unavailable, record `POSITION_CLASS_UNAVAILABLE` and omit this family without replacing it. ## Frozen cell support rule A cell is `SUPPORTED` only when it contains at least `25,000` DEVELOPMENT candidate rows and at least `500` complete-game groups. Unsupported cells remain in the artifact table but have no routing authority. ## Frozen metrics For each supported cell and each of the three fixed prediction streams, report: - rows and groups; - RMSE; - MAE; - signed bias; - correlation/R2 where numerically defined; - the same metrics in each of the four inherited grouped folds. For each additive candidate versus accepted baseline, additionally report: - RMSE regression (`candidate_rmse - baseline_rmse`); - MAE regression (`candidate_mae - baseline_mae`); - squared-error regression mass, defined as the sum over rows of `candidate_squared_error - baseline_squared_error`; - positive squared-error regression mass, defined by clipping row-level regression at zero before summing; - within-family share of positive squared-error regression mass; - count of folds with positive RMSE regression; - count of folds with positive MAE regression. No metric other than the frozen metrics may determine routing. ## Prospectively frozen routing rule Evaluate routing primarily on `native-cubeful-p3-context-additive-ridge-v1`, because it is the Generation 1 candidate that explicitly added the full accepted factual context block. Use P3-only results as a descriptive control. A supported cell is `STABLE_CONTEXT_REGRESSION` only if: 1. RMSE regression is at least `+0.005000` globally within the cell; 2. MAE regression is strictly positive; 3. RMSE regression is positive in all `4/4` inherited grouped folds; 4. MAE regression is positive in at least `3/4` inherited grouped folds. Classify `CONTEXT_FAILURE_LOCALIZED` if at least one `STABLE_CONTEXT_REGRESSION` cell has within-family positive squared-error regression mass share `>= 0.35`. Select exactly one routing cell by highest share, then largest RMSE regression, then lexicographic family/cell name as deterministic tie-breakers. Otherwise classify `CONTEXT_FAILURE_DISTRIBUTED` if at least three `STABLE_CONTEXT_REGRESSION` cells exist across at least two different segment families. Otherwise classify `NO_STABLE_CONTEXT_FAILURE_STRUCTURE`. The accepted baseline's highest-error supported cells are descriptive evidence only and cannot change the terminal routing classification. ## Follow-up authority If `CONTEXT_FAILURE_LOCALIZED`, the next separately frozen experiment may test one bounded context-family representation or ablation directly tied to the single routing cell/family. It may not search multiple families. If `CONTEXT_FAILURE_DISTRIBUTED`, do not fit another local context patch. The next separately frozen question should test whether the additive native-Cubeful failure is a broader target/model-capacity mismatch using a bounded non-context-specific hypothesis. If `NO_STABLE_CONTEXT_FAILURE_STRUCTURE`, close the current factual-context patch line unless another completed experiment provides a new causal hypothesis. No outcome in this experiment authorizes calculated cubeful logic, a production change, or protected evaluation. ## Required durable artifacts Record and hash: - input/prediction/partition receipt; - raw-to-semantic field mapping receipt; - exact segment-cell membership summary; - complete cell metrics for all three streams; - four-fold evidence; - additive-versus-baseline regression-mass tables; - support and stability decisions; - deterministic terminal routing record; - protected-access log showing exactly `0`; - deterministic verification record; - package manifest and SHA256SUMS. At completion update the lane result and current handoff, commit implementation and Task Management changes locally, and leave the lane `COMPLETED_WAITING_FOR_RESEARCH_DIRECTOR`. The supervisor owns pushing the clean branches. ## HFCS safety Before substantive computation perform fresh CPU, RAM, disk, and process preflight. Run Explainer work niced with conservative numerical threading. Corpus Listener/writer, Post Match, and existing GNU/Sage processes retain priority and must not be signaled, reniced, stopped, restarted, reconfigured, or otherwise mutated. If substantial headroom is not available, stop this lane with a durable host-safety blocker rather than interfering with priority work. ## Absolute boundaries Existing accepted data only. New GNU `0`; new Sage `0`; new matches `0`; new labels `0`; generic new 0-ply `0`; Sage/GNU campaign validation data as training `NO`; production promotion `NO`; Analyzer mutation `NO`; Canonical mutation `NO`; Corpus mutation `NO`; calculated-cubeful authority change `NO`. Accepted Ridge and HADD product roles remain unchanged. Any future improved research model is only `CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION`. codex The frozen chain is now clear: Generation 3 permits one—and only one—fit, removing `cubeful_crawford` from the Generation 1 additive context model. Before any model work, I’m performing the required non-mutating HFCS safety preflight and checking both implementation and Task Management histories against the frozen starting authorities. exec /bin/bash -lc "uptime free -h df -h /users/a2andrad/scratch /tmp ps -eo pid,ppid,ni,stat,%cpu,%mem,comm,args --sort=-%cpu | sed -n '1,80p' git status --short --branch git rev-parse HEAD git log -5 --oneline --decorate" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 142ms: 04:10:41 up 42 days, 6:54, 15 users, load average: 4.26, 4.28, 4.46 total used free shared buff/cache available Mem: 157Gi 56Gi 19Gi 1.8Gi 85Gi 100Gi Swap: 0B 0B 0B Filesystem Size Used Avail Use% Mounted on ranch-storage.csclub.uwaterloo.ca:/users 13T 5.8T 7.2T 45% /users tmpfs 79G 1.9G 77G 3% /tmp PID PPID NI STAT %CPU %MEM COMMAND COMMAND 1 0 19 SNs 58.3 0.0 codex codex-linux-sandbox --sandbox-policy-cwd /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl --command-cwd /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl --permission-profile {"type":"managed","file_system":{"type":"restricted","entries":[{"path":{"type":"special","value":{"kind":"root"}},"access":"read"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl"},"access":"write"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm"},"access":"write"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.git"},"access":"write"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.git"},"access":"write"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001"},"access":"write"},{"path":{"type":"special","value":{"kind":"slash_tmp"}},"access":"write"},{"path":{"type":"special","value":{"kind":"tmpdir"}},"access":"write"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.git"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.git"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.git/.git"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.git/.git"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/.git"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.agents"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.agents"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.git/.agents"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.git/.agents"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/.agents"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.codex"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.codex"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.git/.codex"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.git/.codex"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/.codex"},"access":"read","missing_path_behavior":"skip"}]},"network":"restricted"} --apply-seccomp-then-exec -- /bin/bash -c __CODEX_SNAPSHOT_OVERRIDE_SET_0="${CODEX_APPLY_PATCH_PRESERVE_LINE_ENDINGS+x}" __CODEX_SNAPSHOT_OVERRIDE_0="${CODEX_APPLY_PATCH_PRESERVE_LINE_ENDINGS-}" __CODEX_SNAPSHOT_OVERRIDE_SET_1="${CODEX_PERMISSION_PROFILE+x}" __CODEX_SNAPSHOT_OVERRIDE_1="${CODEX_PERMISSION_PROFILE-}" __CODEX_SNAPSHOT_OVERRIDE_SET_2="${CODEX_PLUGIN_METRICS_OUTPUT+x}" __CODEX_SNAPSHOT_OVERRIDE_2="${CODEX_PLUGIN_METRICS_OUTPUT-}" __CODEX_SNAPSHOT_OVERRIDE_SET_3="${CODEX_SESSION_ID+x}" __CODEX_SNAPSHOT_OVERRIDE_3="${CODEX_SESSION_ID-}" __CODEX_SNAPSHOT_OVERRIDE_SET_4="${CODEX_THREAD_ID+x}" __CODEX_SNAPSHOT_OVERRIDE_4="${CODEX_THREAD_ID-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_0="${ALL_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_0="${ALL_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_1="${BUNDLE_HTTPS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_1="${BUNDLE_HTTPS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_2="${BUNDLE_HTTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_2="${BUNDLE_HTTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_3="${BUNDLE_NO_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_3="${BUNDLE_NO_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_4="${BUNDLE_SSL_CA_CERT+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_4="${BUNDLE_SSL_CA_CERT-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_5="${CARGO_HTTP_CAINFO+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_5="${CARGO_HTTP_CAINFO-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_6="${CODEX_CA_CERTIFICATE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_6="${CODEX_CA_CERTIFICATE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_7="${CODEX_NETWORK_ALLOW_LOCAL_BINDING+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_7="${CODEX_NETWORK_ALLOW_LOCAL_BINDING-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_8="${CODEX_NETWORK_PROXY_ACTIVE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_8="${CODEX_NETWORK_PROXY_ACTIVE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_9="${CODEX_NETWORK_PROXY_ATTRIBUTION+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_9="${CODEX_NETWORK_PROXY_ATTRIBUTION-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_10="${CODEX_NETWORK_PROXY_BROKERED_CREDENTIALS+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_10="${CODEX_NETWORK_PROXY_BROKERED_CREDENTIALS-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_11="${CODEX_NETWORK_PROXY_CREDENTIAL_BROKER_ACTIVE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_11="${CODEX_NETWORK_PROXY_CREDENTIAL_BROKER_ACTIVE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_12="${CURL_CA_BUNDLE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_12="${CURL_CA_BUNDLE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_13="${DOCKER_HTTPS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_13="${DOCKER_HTTPS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_14="${DOCKER_HTTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_14="${DOCKER_HTTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_15="${ELECTRON_GET_USE_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_15="${ELECTRON_GET_USE_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_16="${FTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_16="${FTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_17="${GIT_SSL_CAINFO+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_17="${GIT_SSL_CAINFO-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_18="${HTTPS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_18="${HTTPS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_19="${HTTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_19="${HTTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_20="${NODE_EXTRA_CA_CERTS+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_20="${NODE_EXTRA_CA_CERTS-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_21="${NODE_USE_ENV_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_21="${NODE_USE_ENV_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_22="${NO_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_22="${NO_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_23="${NPM_CONFIG_CAFILE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_23="${NPM_CONFIG_CAFILE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_24="${NPM_CONFIG_HTTPS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_24="${NPM_CONFIG_HTTPS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_25="${NPM_CONFIG_HTTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_25="${NPM_CONFIG_HTTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_26="${NPM_CONFIG_NOPROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_26="${NPM_CONFIG_NOPROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_27="${NPM_CONFIG_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_27="${NPM_CONFIG_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_28="${PIP_CERT+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_28="${PIP_CERT-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_29="${PIP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_29="${PIP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_30="${REQUESTS_CA_BUNDLE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_30="${REQUESTS_CA_BUNDLE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_31="${SSL_CERT_FILE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_31="${SSL_CERT_FILE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_32="${WSS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_32="${WSS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_33="${WS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_33="${WS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_34="${YARN_HTTPS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_34="${YARN_HTTPS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_35="${YARN_HTTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_35="${YARN_HTTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_36="${YARN_NO_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_36="${YARN_NO_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_37="${all_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_37="${all_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_38="${ftp_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_38="${ftp_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_39="${http_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_39="${http_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_40="${https_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_40="${https_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_41="${no_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_41="${no_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_42="${npm_config_cafile+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_42="${npm_config_cafile-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_43="${npm_config_http_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_43="${npm_config_http_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_44="${npm_config_https_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_44="${npm_config_https_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_45="${npm_config_noproxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_45="${npm_config_noproxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_46="${npm_config_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_46="${npm_config_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_47="${ws_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_47="${ws_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_48="${wss_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_48="${wss_proxy-}" __CODEX_SNAPSHOT_PROXY_ENV_SET="${CODEX_NETWORK_PROXY_ACTIVE+x}" if . '/users/a2andrad/.codex/shell_snapshots/01a032d1-af70-7903-99d2-14de6b7620fe.1787558997909339012.sh' >/dev/null 2>&1; then :; fi if [ -n "${__CODEX_SNAPSHOT_OVERRIDE_SET_0}" ]; then export CODEX_APPLY_PATCH_PRESERVE_LINE_ENDINGS="${__CODEX_SNAPSHOT_OVERRIDE_0}"; else unset CODEX_APPLY_PATCH_PRESERVE_LINE_ENDINGS; fi if [ -n "${__CODEX_SNAPSHOT_OVERRIDE_SET_1}" ]; then export CODEX_PERMISSION_PROFILE="${__CODEX_SNAPSHOT_OVERRIDE_1}"; else unset CODEX_PERMISSION_PROFILE; fi if [ -n "${__CODEX_SNAPSHOT_OVERRIDE_SET_2}" ]; then export CODEX_PLUGIN_METRICS_OUTPUT="${__CODEX_SNAPSHOT_OVERRIDE_2}"; else unset CODEX_PLUGIN_METRICS_OUTPUT; fi if [ -n "${__CODEX_SNAPSHOT_OVERRIDE_SET_3}" ]; then export CODEX_SESSION_ID="${__CODEX_SNAPSHOT_OVERRIDE_3}"; else unset CODEX_SESSION_ID; fi if [ -n "${__CODEX_SNAPSHOT_OVERRIDE_SET_4}" ]; then export CODEX_THREAD_ID="${__CODEX_SNAPSHOT_OVERRIDE_4}"; else unset CODEX_THREAD_ID; fi if [ -n "$__CODEX_SNAPSHOT_PROXY_ENV_SET" ] || [ -n "${CODEX_NETWORK_PROXY_ACTIVE+x}" ]; then if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_0}" ]; then export ALL_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_0}"; else unset ALL_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_1}" ]; then export BUNDLE_HTTPS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_1}"; else unset BUNDLE_HTTPS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_2}" ]; then export BUNDLE_HTTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_2}"; else unset BUNDLE_HTTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_3}" ]; then export BUNDLE_NO_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_3}"; else unset BUNDLE_NO_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_4}" ]; then export BUNDLE_SSL_CA_CERT="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_4}"; else unset BUNDLE_SSL_CA_CERT; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_5}" ]; then export CARGO_HTTP_CAINFO="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_5}"; else unset CARGO_HTTP_CAINFO; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_6}" ]; then export CODEX_CA_CERTIFICATE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_6}"; else unset CODEX_CA_CERTIFICATE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_7}" ]; then export CODEX_NETWORK_ALLOW_LOCAL_BINDING="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_7}"; else unset CODEX_NETWORK_ALLOW_LOCAL_BINDING; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_8}" ]; then export CODEX_NETWORK_PROXY_ACTIVE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_8}"; else unset CODEX_NETWORK_PROXY_ACTIVE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_9}" ]; then export CODEX_NETWORK_PROXY_ATTRIBUTION="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_9}"; else unset CODEX_NETWORK_PROXY_ATTRIBUTION; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_10}" ]; then export CODEX_NETWORK_PROXY_BROKERED_CREDENTIALS="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_10}"; else unset CODEX_NETWORK_PROXY_BROKERED_CREDENTIALS; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_11}" ]; then export CODEX_NETWORK_PROXY_CREDENTIAL_BROKER_ACTIVE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_11}"; else unset CODEX_NETWORK_PROXY_CREDENTIAL_BROKER_ACTIVE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_12}" ]; then export CURL_CA_BUNDLE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_12}"; else unset CURL_CA_BUNDLE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_13}" ]; then export DOCKER_HTTPS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_13}"; else unset DOCKER_HTTPS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_14}" ]; then export DOCKER_HTTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_14}"; else unset DOCKER_HTTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_15}" ]; then export ELECTRON_GET_USE_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_15}"; else unset ELECTRON_GET_USE_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_16}" ]; then export FTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_16}"; else unset FTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_17}" ]; then export GIT_SSL_CAINFO="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_17}"; else unset GIT_SSL_CAINFO; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_18}" ]; then export HTTPS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_18}"; else unset HTTPS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_19}" ]; then export HTTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_19}"; else unset HTTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_20}" ]; then export NODE_EXTRA_CA_CERTS="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_20}"; else unset NODE_EXTRA_CA_CERTS; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_21}" ]; then export NODE_USE_ENV_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_21}"; else unset NODE_USE_ENV_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_22}" ]; then export NO_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_22}"; else unset NO_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_23}" ]; then export NPM_CONFIG_CAFILE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_23}"; else unset NPM_CONFIG_CAFILE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_24}" ]; then export NPM_CONFIG_HTTPS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_24}"; else unset NPM_CONFIG_HTTPS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_25}" ]; then export NPM_CONFIG_HTTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_25}"; else unset NPM_CONFIG_HTTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_26}" ]; then export NPM_CONFIG_NOPROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_26}"; else unset NPM_CONFIG_NOPROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_27}" ]; then export NPM_CONFIG_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_27}"; else unset NPM_CONFIG_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_28}" ]; then export PIP_CERT="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_28}"; else unset PIP_CERT; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_29}" ]; then export PIP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_29}"; else unset PIP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_30}" ]; then export REQUESTS_CA_BUNDLE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_30}"; else unset REQUESTS_CA_BUNDLE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_31}" ]; then export SSL_CERT_FILE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_31}"; else unset SSL_CERT_FILE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_32}" ]; then export WSS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_32}"; else unset WSS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_33}" ]; then export WS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_33}"; else unset WS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_34}" ]; then export YARN_HTTPS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_34}"; else unset YARN_HTTPS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_35}" ]; then export YARN_HTTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_35}"; else unset YARN_HTTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_36}" ]; then export YARN_NO_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_36}"; else unset YARN_NO_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_37}" ]; then export all_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_37}"; else unset all_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_38}" ]; then export ftp_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_38}"; else unset ftp_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_39}" ]; then export http_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_39}"; else unset http_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_40}" ]; then export https_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_40}"; else unset https_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_41}" ]; then export no_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_41}"; else unset no_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_42}" ]; then export npm_config_cafile="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_42}"; else unset npm_config_cafile; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_43}" ]; then export npm_config_http_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_43}"; else unset npm_config_http_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_44}" ]; then export npm_config_https_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_44}"; else unset npm_config_https_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_45}" ]; then export npm_config_noproxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_45}"; else unset npm_config_noproxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_46}" ]; then export npm_config_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_46}"; else unset npm_config_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_47}" ]; then export ws_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_47}"; else unset ws_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_48}" ]; then export wss_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_48}"; else unset wss_proxy; fi fi if [ -n "${PATH:-}" ]; then export PATH='/users/a2andrad/.codex/packages/standalone/releases/0.149.0-x86_64-unknown-linux-musl/codex-path':"$PATH"; else export PATH='/users/a2andrad/.codex/packages/standalone/releases/0.149.0-x86_64-unknown-linux-musl/codex-path'; fi exec '/bin/bash' -c 'uptime free -h df -h /users/a2andrad/scratch /tmp ps -eo pid,ppid,ni,stat,%cpu,%mem,comm,args --sort=-%cpu | sed -n '"'"'1,80p'"'"' git status --short --branch git rev-parse HEAD git log -5 --oneline --decorate' 2 1 19 SN 25.0 0.0 bash /bin/bash -c uptime free -h df -h /users/a2andrad/scratch /tmp ps -eo pid,ppid,ni,stat,%cpu,%mem,comm,args --sort=-%cpu | sed -n '1,80p' git status --short --branch git rev-parse HEAD git log -5 --oneline --decorate 6 2 19 RN 0.0 0.0 ps ps -eo pid,ppid,ni,stat,%cpu,%mem,comm,args --sort=-%cpu 7 2 19 SN 0.0 0.0 sed sed -n 1,80p ## research/explainer-native-cubeful-k001...origin/research/explainer-native-cubeful-k001 75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a 75ce3ba (HEAD -> research/explainer-native-cubeful-k001, origin/research/explainer-native-cubeful-k001) Complete native Cubeful context error diagnosis 2b2a822 Complete additive native Cubeful experiment 18135ef Freeze native Cubeful development decision 8eacfc0 Freeze native Cubeful candidate models cb9534f Freeze additive native Cubeful experiment exec /bin/bash -lc 'git status --short --branch git rev-parse HEAD git log -6 --oneline --decorate git show --stat --oneline 222c2a50418082c65fd57792564c1dc1cc1ad0f6 --' in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm succeeded in 131ms: ## milestone/explainer-native-cubeful-k001...origin/milestone/explainer-native-cubeful-k001 9e234525ebd87ff115f8532bcb8b4d2cf888b441 9e23452 (HEAD -> milestone/explainer-native-cubeful-k001, origin/milestone/explainer-native-cubeful-k001) Route native cubeful Generation 3 Crawford-state ablation 222c2a5 Freeze native cubeful Generation 3 Crawford-state ablation d0521ba Complete native cubeful Generation 2 handoff 9b0d87f Route native cubeful lane to frozen Generation 2 b4c08ba Freeze native cubeful Generation 2 context error diagnosis 0e3a955 Complete native Cubeful research handoff 222c2a5 Freeze native cubeful Generation 3 Crawford-state ablation ...st-crawford-state-context-family-ablation-v1.md | 189 +++++++++++++++++++++ 1 file changed, 189 insertions(+) → Verify frozen authorities, host safety, repositories, and accepted input identities • Implement and test the single Crawford-field ablation runner without protected access • Run the one frozen fit and produce hashed durable results • Verify artifacts, update handoff/result, and commit both repositories locally exec /bin/bash -lc "git show --stat --summary 2b2a82284649485b00621cd243dfa17c6accca9c git show --stat --summary 75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a rg --files artifacts/development | rg 'native|cubeful' | sed -n '1,260p' rg -n \"native-cubeful|cubeful_crawford|additive-ridge|streaming Adam|20260823\" scripts src tests artifacts/development results docs/modeling -g '*.py' -g '*.json' -g '*.md' | sed -n '1,320p'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 278ms: commit 2b2a82284649485b00621cd243dfa17c6accca9c Author: Marty Gale <300671762+backgammon-made-simple@users.noreply.github.com> Date: Mon Aug 24 02:56:19 2026 -0400 Complete additive native Cubeful experiment .../SHA256SUMS | 10 + .../manifest.json | 56 +++++ .../preflight-log.jsonl | 1 + .../protected-final.json | 7 + .../result-summary.json | 234 +++++++++++++++++++++ .../self-verification.json | 68 ++++++ .../test-evidence.json | 27 +++ .../improve-additive-native-cubeful-modeling-v1.md | 29 ++- ...mprove-additive-native-cubeful-modeling-v1.json | 234 +++++++++++++++++++++ .../native_cubeful_experiment.py | 6 +- 10 files changed, 666 insertions(+), 6 deletions(-) create mode 100644 artifacts/development/improve-additive-native-cubeful-modeling-v1/SHA256SUMS create mode 100644 artifacts/development/improve-additive-native-cubeful-modeling-v1/manifest.json create mode 100644 artifacts/development/improve-additive-native-cubeful-modeling-v1/protected-final.json create mode 100644 artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json create mode 100644 artifacts/development/improve-additive-native-cubeful-modeling-v1/self-verification.json create mode 100644 artifacts/development/improve-additive-native-cubeful-modeling-v1/test-evidence.json create mode 100644 results/improve-additive-native-cubeful-modeling-v1.json commit 75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a Author: Marty Gale <300671762+backgammon-made-simple@users.noreply.github.com> Date: Mon Aug 24 03:32:20 2026 -0400 Complete native Cubeful context error diagnosis .../SHA256SUMS | 16 + .../cell-metrics.json | 1001 ++++++ .../deterministic-verification.json | 128 + .../field-mapping-receipt.json | 93 + .../fold-evidence.json | 3617 ++++++++++++++++++++ .../generation-1-reproduction.json | 128 + .../input-receipt.json | 117 + .../manifest.json | 86 + .../prediction-stream-identities.json | 18 + .../preflight-log.jsonl | 1 + .../protected-access-log.json | 8 + .../regression-mass.json | 594 ++++ .../result-summary.json | 119 + .../routing-record.json | 26 + .../run-record.json | 21 + .../segment-membership.json | 206 ++ .../self-verification.json | 80 + .../support-stability.json | 577 ++++ .../test-evidence.json | 25 + ...se-native-cubeful-context-error-structure-v1.md | 63 + ...-native-cubeful-context-error-structure-v1.json | 119 + scripts/run_native_cubeful_context_diagnostic.py | 69 + .../native_cubeful_context_diagnostic.py | 975 ++++++ tests/test_native_cubeful_context_diagnostic.py | 76 + 24 files changed, 8163 insertions(+) create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/SHA256SUMS create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/deterministic-verification.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/field-mapping-receipt.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/generation-1-reproduction.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/manifest.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/prediction-stream-identities.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/preflight-log.jsonl create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/protected-access-log.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/regression-mass.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/result-summary.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/routing-record.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/run-record.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/segment-membership.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/self-verification.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/support-stability.json create mode 100644 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/test-evidence.json create mode 100644 docs/modeling/diagnose-native-cubeful-context-error-structure-v1.md create mode 100644 results/diagnose-native-cubeful-context-error-structure-v1.json create mode 100644 scripts/run_native_cubeful_context_diagnostic.py create mode 100644 src/backgammon_explainer/native_cubeful_context_diagnostic.py create mode 100644 tests/test_native_cubeful_context_diagnostic.py artifacts/development/diagnose-native-cubeful-context-error-structure-v1/support-stability.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/regression-mass.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/routing-record.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/deterministic-verification.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/prediction-stream-identities.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/SHA256SUMS artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/field-mapping-receipt.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/protected-access-log.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/run-record.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/segment-membership.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/self-verification.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/manifest.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/preflight-log.jsonl artifacts/development/diagnose-native-cubeful-context-error-structure-v1/test-evidence.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/generation-1-reproduction.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/result-summary.json artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json artifacts/development/explainer-k002-position-value-modeling/direct-cubeful.json artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json artifacts/development/improve-additive-native-cubeful-modeling-v1/development.json artifacts/development/improve-additive-native-cubeful-modeling-v1/winner-freeze.json artifacts/development/improve-additive-native-cubeful-modeling-v1/protected-final.json artifacts/development/improve-additive-native-cubeful-modeling-v1/preflight-log.jsonl artifacts/development/improve-additive-native-cubeful-modeling-v1/test-evidence.json artifacts/development/improve-additive-native-cubeful-modeling-v1/frozen-authorities.json artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json artifacts/development/improve-additive-native-cubeful-modeling-v1/contribution-evidence.json artifacts/development/improve-additive-native-cubeful-modeling-v1/self-verification.json artifacts/development/improve-additive-native-cubeful-modeling-v1/manifest.json artifacts/development/improve-additive-native-cubeful-modeling-v1/models.json artifacts/development/improve-additive-native-cubeful-modeling-v1/protected-access-log.json artifacts/development/improve-additive-native-cubeful-modeling-v1/SHA256SUMS docs/modeling/improve-additive-native-cubeful-modeling-v1.md:29:`config/modeling/improve-additive-native-cubeful-modeling-v1.json` before any results/diagnose-native-cubeful-context-error-structure-v1.json:17: "native-cubeful-p3-additive-ridge-v1": "32a9e5746c5ca19344a437d1cb2e5f85ad86f7ad945d569f43b9617bc93b3a4e", results/diagnose-native-cubeful-context-error-structure-v1.json:18: "native-cubeful-p3-context-additive-ridge-v1": "995b3e1463272131b7ae06d2f0996c4bf8feabe2db72886434fb69ac4b7a4456" results/diagnose-native-cubeful-context-error-structure-v1.json:118: "version": "diagnose-native-cubeful-context-error-structure-v1-result-v1" src/backgammon_explainer/constrained_additive_position_model.py:46:INNER_SEED = 20260823 src/backgammon_explainer/constrained_additive_position_model.py:616: "algorithm": "deterministic streaming Adam", src/backgammon_explainer/constrained_additive_position_model.py:661: "algorithm": "deterministic streaming Adam on Ridge objective", src/backgammon_explainer/constrained_additive_position_model.py:949: "algorithm": "deterministic streaming Adam warm start with Polyak average over final two epochs", docs/modeling/diagnose-native-cubeful-context-error-structure-v1.md:61:`artifacts/development/diagnose-native-cubeful-context-error-structure-v1/`; docs/modeling/diagnose-native-cubeful-context-error-structure-v1.md:63:`results/diagnose-native-cubeful-context-error-structure-v1.json`. src/backgammon_explainer/native_cubeful_context_diagnostic.py:6:in diagnose-native-cubeful-context-error-structure-v1. src/backgammon_explainer/native_cubeful_context_diagnostic.py:48:VERSION = "diagnose-native-cubeful-context-error-structure-v1" src/backgammon_explainer/native_cubeful_context_diagnostic.py:59:PRIMARY_CANDIDATE = "native-cubeful-p3-context-additive-ridge-v1" src/backgammon_explainer/native_cubeful_context_diagnostic.py:186: "raw_fields": ["cubeful_is_money", "cubeful_match_length", "cubeful_crawford"], scripts/run_native_cubeful_context_diagnostic.py:2:"""Phase runner for diagnose-native-cubeful-context-error-structure-v1.""" scripts/run_native_cubeful_context_diagnostic.py:22:PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/002-diagnose-native-cubeful-context-error-structure-v1.md") scripts/run_native_cubeful_context_diagnostic.py:25:GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") scripts/run_native_cubeful_context_diagnostic.py:26:ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") scripts/run_native_cubeful_context_diagnostic.py:27:RESULT = Path("results/diagnose-native-cubeful-context-error-structure-v1.json") results/improve-additive-native-cubeful-modeling-v1.json:55: "native-cubeful-p3-additive-ridge-v1": { results/improve-additive-native-cubeful-modeling-v1.json:89: "native-cubeful-p3-context-additive-ridge-v1": { results/improve-additive-native-cubeful-modeling-v1.json:133: "native-cubeful-p3-additive-ridge-v1": { results/improve-additive-native-cubeful-modeling-v1.json:141: "native-cubeful-p3-context-additive-ridge-v1": { results/improve-additive-native-cubeful-modeling-v1.json:152: "native-cubeful-p3-additive-ridge-v1": "74a62b141204d65ecc70b50534cab260a7cca7ab08d63e4a927e8021e9b48044", results/improve-additive-native-cubeful-modeling-v1.json:153: "native-cubeful-p3-context-additive-ridge-v1": "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" results/improve-additive-native-cubeful-modeling-v1.json:186: "version": "improve-additive-native-cubeful-modeling-v1-protected-v1", results/improve-additive-native-cubeful-modeling-v1.json:199: "version": "improve-additive-native-cubeful-modeling-v1-result-v1", results/improve-additive-native-cubeful-modeling-v1.json:203: "native-cubeful-p3-additive-ridge-v1": { results/improve-additive-native-cubeful-modeling-v1.json:212: "native-cubeful-p3-context-additive-ridge-v1": { results/improve-additive-native-cubeful-modeling-v1.json:229: "version": "improve-additive-native-cubeful-modeling-v1-winner-freeze-v1", tests/test_native_cubeful_experiment.py:22:CONFIG = Path("config/modeling/improve-additive-native-cubeful-modeling-v1.json") tests/test_native_cubeful_experiment.py:39: rng = np.random.default_rng(20260823) tests/test_native_cubeful_experiment.py:57: assert DEVELOPMENT_FOLD_SEED == 20260823 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/support-stability.json:576: "version": "diagnose-native-cubeful-context-error-structure-v1-support-stability-v1" tests/test_position_value_modeling.py:81: rng = np.random.default_rng(20260823) artifacts/development/diagnose-native-cubeful-context-error-structure-v1/regression-mass.json:3: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/regression-mass.json:296: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/regression-mass.json:593: "version": "diagnose-native-cubeful-context-error-structure-v1-regression-mass-v1" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/routing-record.json:6: "primary_candidate": "native-cubeful-p3-context-additive-ridge-v1", artifacts/development/diagnose-native-cubeful-context-error-structure-v1/routing-record.json:25: "version": "diagnose-native-cubeful-context-error-structure-v1-routing-v1" tests/test_constrained_additive_position_model.py:22: rng = np.random.default_rng(20260823) artifacts/development/diagnose-native-cubeful-context-error-structure-v1/deterministic-verification.json:127: "version": "diagnose-native-cubeful-context-error-structure-v1-deterministic-verification-v1" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/prediction-stream-identities.json:4: "native-cubeful-p3-additive-ridge-v1": "6eca0be6638918d54dab8c464ab5271010e8884a6c1441f4476a3befc6c0d345", artifacts/development/diagnose-native-cubeful-context-error-structure-v1/prediction-stream-identities.json:5: "native-cubeful-p3-context-additive-ridge-v1": "8c1d1454054393b8dd4b0902d6d44da0f663a3bd8d7f5a7154181b8dc8e3197a" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/prediction-stream-identities.json:11: "native-cubeful-p3-additive-ridge-v1": "32a9e5746c5ca19344a437d1cb2e5f85ad86f7ad945d569f43b9617bc93b3a4e", artifacts/development/diagnose-native-cubeful-context-error-structure-v1/prediction-stream-identities.json:12: "native-cubeful-p3-context-additive-ridge-v1": "995b3e1463272131b7ae06d2f0996c4bf8feabe2db72886434fb69ac4b7a4456" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/prediction-stream-identities.json:17: "version": "diagnose-native-cubeful-context-error-structure-v1-prediction-stream-identities-v1" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:88: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:97: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:117: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:126: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:146: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:155: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:175: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:184: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:204: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:213: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:235: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:244: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:264: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:273: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:293: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:302: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:322: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:331: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:353: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:362: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:382: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:391: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:411: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:420: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:440: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:449: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:469: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:478: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:498: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:507: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:527: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:536: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:558: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:567: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:587: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:596: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:616: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:625: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:645: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:654: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:674: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:683: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:703: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:712: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:734: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:743: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:763: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:772: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:792: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:801: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:821: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:830: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:852: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:861: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:881: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:890: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:910: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:919: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:939: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:948: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:968: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:977: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json:1000: "version": "diagnose-native-cubeful-context-error-structure-v1-cell-metrics-v1" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/field-mapping-receipt.json:17: "cubeful_crawford" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/field-mapping-receipt.json:92: "version": "diagnose-native-cubeful-context-error-structure-v1-field-mapping-v1" src/backgammon_explainer/native_cubeful_experiment.py:46:VERSION = "improve-additive-native-cubeful-modeling-v1" src/backgammon_explainer/native_cubeful_experiment.py:55:DEVELOPMENT_FOLD_VERSION = "native-cubeful-development-complete-game-fold-4-v1" src/backgammon_explainer/native_cubeful_experiment.py:56:DEVELOPMENT_FOLD_SEED = 20260823 src/backgammon_explainer/native_cubeful_experiment.py:58: "native-cubeful-p3-additive-ridge-v1", src/backgammon_explainer/native_cubeful_experiment.py:59: "native-cubeful-p3-context-additive-ridge-v1", src/backgammon_explainer/native_cubeful_experiment.py:415: "algorithm": "deterministic streaming Adam on Ridge objective", src/backgammon_explainer/native_cubeful_experiment.py:419: "beta1": 0.9, "beta2": 0.999, "seed": 20260823, artifacts/development/diagnose-native-cubeful-context-error-structure-v1/protected-access-log.json:7: "version": "diagnose-native-cubeful-context-error-structure-v1-protected-access-log-v1" src/backgammon_explainer/position_value_modeling.py:161: PositionFeatureDefinition("cubeful_crawford", "CUBEFUL_CONTEXT", "match_context", "boolean", "One for a source-supported Crawford game; zero in money play.", source_authority="gnu_match_id_native"), artifacts/development/diagnose-native-cubeful-context-error-structure-v1/run-record.json:20: "version": "diagnose-native-cubeful-context-error-structure-v1-run-record-v1" scripts/run_native_cubeful_experiment.py:2:"""Phase runner for improve-additive-native-cubeful-modeling-v1.""" scripts/run_native_cubeful_experiment.py:26:CONFIG = Path("config/modeling/improve-additive-native-cubeful-modeling-v1.json") scripts/run_native_cubeful_experiment.py:29:ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") scripts/run_native_cubeful_experiment.py:30:RUNTIME = Path("../runtime/native-cubeful-training-cache") scripts/run_native_cubeful_experiment.py:84: result_path=Path("results/improve-additive-native-cubeful-modeling-v1.json"), artifacts/development/diagnose-native-cubeful-context-error-structure-v1/segment-membership.json:205: "version": "diagnose-native-cubeful-context-error-structure-v1-segment-membership-v1" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/self-verification.json:79: "version": "diagnose-native-cubeful-context-error-structure-v1-self-verification-v1" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/generation-1-reproduction.json:46: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/generation-1-reproduction.json:86: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/generation-1-reproduction.json:127: "version": "diagnose-native-cubeful-context-error-structure-v1-generation-1-reproduction-v1" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/manifest.json:85: "version": "diagnose-native-cubeful-context-error-structure-v1-manifest-v1" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/result-summary.json:17: "native-cubeful-p3-additive-ridge-v1": "32a9e5746c5ca19344a437d1cb2e5f85ad86f7ad945d569f43b9617bc93b3a4e", artifacts/development/diagnose-native-cubeful-context-error-structure-v1/result-summary.json:18: "native-cubeful-p3-context-additive-ridge-v1": "995b3e1463272131b7ae06d2f0996c4bf8feabe2db72886434fb69ac4b7a4456" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/result-summary.json:118: "version": "diagnose-native-cubeful-context-error-structure-v1-result-v1" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/test-evidence.json:24: "version": "diagnose-native-cubeful-context-error-structure-v1-test-evidence-v1" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json:36: "seed": 20260823, artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json:37: "version": "native-cubeful-development-complete-game-fold-4-v1" artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json:75: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json:79: "model_id": "native-cubeful-p3-additive-ridge-v1", artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json:84: "stream": "native-cubeful-p3-additive-ridge-v1", artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json:87: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json:91: "model_id": "native-cubeful-p3-context-additive-ridge-v1", artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json:96: "stream": "native-cubeful-p3-context-additive-ridge-v1", artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json:102: "path": "../tm/milestones/explainer-native-cubeful-k001/prompts/002-diagnose-native-cubeful-context-error-structure-v1.md", artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json:116: "version": "diagnose-native-cubeful-context-error-structure-v1-input-receipt-v1" artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json:55: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json:89: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json:133: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json:141: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json:152: "native-cubeful-p3-additive-ridge-v1": "74a62b141204d65ecc70b50534cab260a7cca7ab08d63e4a927e8021e9b48044", artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json:153: "native-cubeful-p3-context-additive-ridge-v1": "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json:186: "version": "improve-additive-native-cubeful-modeling-v1-protected-v1", artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json:199: "version": "improve-additive-native-cubeful-modeling-v1-result-v1", artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json:203: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json:212: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json:229: "version": "improve-additive-native-cubeful-modeling-v1-winner-freeze-v1", artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:43: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:81: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:159: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:197: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:275: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:313: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:391: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:429: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:507: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:545: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:625: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:663: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:741: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:779: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:857: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:895: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:973: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1011: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1091: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1129: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1207: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1245: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1323: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1361: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1439: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1477: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1555: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1593: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1671: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1709: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1787: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1825: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1905: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:1943: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2021: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2059: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2137: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2175: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2253: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2291: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2369: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2407: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2485: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2523: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2603: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2641: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2719: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2757: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2835: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2873: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2951: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:2989: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3069: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3107: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3185: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3223: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3301: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3339: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3417: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3455: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3533: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3571: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3612: "fold_seed": 20260823, artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3613: "fold_version": "native-cubeful-development-complete-game-fold-4-v1", artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json:3616: "version": "diagnose-native-cubeful-context-error-structure-v1-fold-evidence-v1" artifacts/development/explainer-k002-position-value-modeling/registries.json:3639: "feature_id": "cubeful_crawford", artifacts/development/improve-additive-native-cubeful-modeling-v1/development.json:14: "fold_seed": 20260823, artifacts/development/improve-additive-native-cubeful-modeling-v1/development.json:15: "fold_version": "native-cubeful-development-complete-game-fold-4-v1", artifacts/development/improve-additive-native-cubeful-modeling-v1/development.json:216: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/improve-additive-native-cubeful-modeling-v1/development.json:414: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/improve-additive-native-cubeful-modeling-v1/development.json:616: "version": "improve-additive-native-cubeful-modeling-v1-development-v1" artifacts/development/improve-additive-native-cubeful-modeling-v1/winner-freeze.json:4: "native-cubeful-p3-additive-ridge-v1": { artifacts/development/improve-additive-native-cubeful-modeling-v1/winner-freeze.json:13: "native-cubeful-p3-context-additive-ridge-v1": { artifacts/development/improve-additive-native-cubeful-modeling-v1/winner-freeze.json:30: "version": "improve-additive-native-cubeful-modeling-v1-winner-freeze-v1", artifacts/development/improve-additive-native-cubeful-modeling-v1/protected-final.json:5: "version": "improve-additive-native-cubeful-modeling-v1-protected-v1", artifacts/development/improve-additive-native-cubeful-modeling-v1/self-verification.json:67: "version": "improve-additive-native-cubeful-modeling-v1-self-verification-v1" artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:8: "cache_root": "/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/native-cubeful-training-cache", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:17: "cache": "../runtime/native-cubeful-training-cache/partition-000", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:25: "cache": "../runtime/native-cubeful-training-cache/partition-001", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:33: "cache": "../runtime/native-cubeful-training-cache/partition-002", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:41: "cache": "../runtime/native-cubeful-training-cache/partition-003", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:49: "cache": "../runtime/native-cubeful-training-cache/partition-004", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:57: "cache": "../runtime/native-cubeful-training-cache/partition-005", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:65: "cache": "../runtime/native-cubeful-training-cache/partition-006", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:73: "cache": "../runtime/native-cubeful-training-cache/partition-007", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:81: "cache": "../runtime/native-cubeful-training-cache/partition-008", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:89: "cache": "../runtime/native-cubeful-training-cache/partition-009", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:97: "cache": "../runtime/native-cubeful-training-cache/partition-010", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:105: "cache": "../runtime/native-cubeful-training-cache/partition-011", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:113: "cache": "../runtime/native-cubeful-training-cache/partition-012", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:121: "cache": "../runtime/native-cubeful-training-cache/partition-013", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:129: "cache": "../runtime/native-cubeful-training-cache/partition-014", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:137: "cache": "../runtime/native-cubeful-training-cache/partition-015", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:145: "cache": "../runtime/native-cubeful-training-cache/partition-016", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:153: "cache": "../runtime/native-cubeful-training-cache/partition-017", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:161: "cache": "../runtime/native-cubeful-training-cache/partition-018", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:169: "cache": "../runtime/native-cubeful-training-cache/partition-019", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:177: "cache": "../runtime/native-cubeful-training-cache/partition-020", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:185: "cache": "../runtime/native-cubeful-training-cache/partition-021", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:193: "cache": "../runtime/native-cubeful-training-cache/partition-022", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:201: "cache": "../runtime/native-cubeful-training-cache/partition-023", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:209: "cache": "../runtime/native-cubeful-training-cache/partition-024", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:217: "cache": "../runtime/native-cubeful-training-cache/partition-025", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:225: "cache": "../runtime/native-cubeful-training-cache/partition-026", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:233: "cache": "../runtime/native-cubeful-training-cache/partition-027", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:241: "cache": "../runtime/native-cubeful-training-cache/partition-028", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:249: "cache": "../runtime/native-cubeful-training-cache/partition-029", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:257: "cache": "../runtime/native-cubeful-training-cache/partition-030", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:265: "cache": "../runtime/native-cubeful-training-cache/partition-031", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:273: "cache": "../runtime/native-cubeful-training-cache/partition-032", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:281: "cache": "../runtime/native-cubeful-training-cache/partition-033", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:289: "cache": "../runtime/native-cubeful-training-cache/partition-034", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:297: "cache": "../runtime/native-cubeful-training-cache/partition-035", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:305: "cache": "../runtime/native-cubeful-training-cache/partition-036", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:313: "cache": "../runtime/native-cubeful-training-cache/partition-037", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:321: "cache": "../runtime/native-cubeful-training-cache/partition-038", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:329: "cache": "../runtime/native-cubeful-training-cache/partition-039", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:337: "cache": "../runtime/native-cubeful-training-cache/partition-040", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:345: "cache": "../runtime/native-cubeful-training-cache/partition-041", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:353: "cache": "../runtime/native-cubeful-training-cache/partition-042", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:361: "cache": "../runtime/native-cubeful-training-cache/partition-043", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:369: "cache": "../runtime/native-cubeful-training-cache/partition-044", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:377: "cache": "../runtime/native-cubeful-training-cache/partition-045", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:385: "cache": "../runtime/native-cubeful-training-cache/partition-046", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:393: "cache": "../runtime/native-cubeful-training-cache/partition-047", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:401: "cache": "../runtime/native-cubeful-training-cache/partition-048", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:409: "cache": "../runtime/native-cubeful-training-cache/partition-049", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:417: "cache": "../runtime/native-cubeful-training-cache/partition-050", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:425: "cache": "../runtime/native-cubeful-training-cache/partition-051", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:433: "cache": "../runtime/native-cubeful-training-cache/partition-052", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:441: "cache": "../runtime/native-cubeful-training-cache/partition-053", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:449: "cache": "../runtime/native-cubeful-training-cache/partition-054", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:457: "cache": "../runtime/native-cubeful-training-cache/partition-055", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:465: "cache": "../runtime/native-cubeful-training-cache/partition-056", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:473: "cache": "../runtime/native-cubeful-training-cache/partition-057", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:481: "cache": "../runtime/native-cubeful-training-cache/partition-058", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:489: "cache": "../runtime/native-cubeful-training-cache/partition-059", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:497: "cache": "../runtime/native-cubeful-training-cache/partition-060", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:505: "cache": "../runtime/native-cubeful-training-cache/partition-061", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:513: "cache": "../runtime/native-cubeful-training-cache/partition-062", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:521: "cache": "../runtime/native-cubeful-training-cache/partition-063", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:529: "cache": "../runtime/native-cubeful-training-cache/partition-064", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:537: "cache": "../runtime/native-cubeful-training-cache/partition-065", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:545: "cache": "../runtime/native-cubeful-training-cache/partition-066", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:553: "cache": "../runtime/native-cubeful-training-cache/partition-067", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:561: "cache": "../runtime/native-cubeful-training-cache/partition-068", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:569: "cache": "../runtime/native-cubeful-training-cache/partition-069", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:577: "cache": "../runtime/native-cubeful-training-cache/partition-070", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:585: "cache": "../runtime/native-cubeful-training-cache/partition-071", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:593: "cache": "../runtime/native-cubeful-training-cache/partition-072", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:601: "cache": "../runtime/native-cubeful-training-cache/partition-073", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:609: "cache": "../runtime/native-cubeful-training-cache/partition-074", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:617: "cache": "../runtime/native-cubeful-training-cache/partition-075", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:625: "cache": "../runtime/native-cubeful-training-cache/partition-076", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:633: "cache": "../runtime/native-cubeful-training-cache/partition-077", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:641: "cache": "../runtime/native-cubeful-training-cache/partition-078", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:649: "cache": "../runtime/native-cubeful-training-cache/partition-079", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:657: "cache": "../runtime/native-cubeful-training-cache/partition-080", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:665: "cache": "../runtime/native-cubeful-training-cache/partition-081", artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json:675: "version": "improve-additive-native-cubeful-modeling-v1-training-cache-v1" artifacts/development/improve-additive-native-cubeful-modeling-v1/test-evidence.json:26: "version": "improve-additive-native-cubeful-modeling-v1-test-evidence-v1" artifacts/development/improve-additive-native-cubeful-modeling-v1/models.json:1417: "model_id": "native-cubeful-p3-additive-ridge-v1", exec /bin/bash -lc "wc -l src/backgammon_explainer/native_cubeful_experiment.py scripts/run_native_cubeful_experiment.py tests/test_native_cubeful_experiment.py config/modeling/improve-additive-native-cubeful-modeling-v1.json sed -n '1,260p' scripts/run_native_cubeful_experiment.py sed -n '1,220p' tests/test_native_cubeful_experiment.py sed -n '1,360p' config/modeling/improve-additive-native-cubeful-modeling-v1.json" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 74ms: 853 src/backgammon_explainer/native_cubeful_experiment.py 95 scripts/run_native_cubeful_experiment.py 85 tests/test_native_cubeful_experiment.py 164 config/modeling/improve-additive-native-cubeful-modeling-v1.json 1197 total #!/usr/bin/env python3 """Phase runner for improve-additive-native-cubeful-modeling-v1.""" from __future__ import annotations import argparse import json from pathlib import Path from backgammon_explainer.native_cubeful_experiment import ( build_manifest, build_summary, build_training_cache, fit_models, freeze_authorities, freeze_winner, record_preflight, score_development, score_protected, verify_package, ) SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") CANONICAL = Path("/users/a2andrad/code/artifacts/explainer-k001/canonical-analysis-v1/canonical-analysis-reference-2c828e118b6cf22f") CONFIG = Path("config/modeling/improve-additive-native-cubeful-modeling-v1.json") SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") RUNTIME = Path("../runtime/native-cubeful-training-cache") def main() -> int: parser = argparse.ArgumentParser() parser.add_argument("phase", choices=( "freeze", "cache", "fit", "score-development", "freeze-winner", "score-protected", "summarize", "manifest", "verify", )) parser.add_argument("--root", type=Path, default=ROOT) parser.add_argument("--cache-root", type=Path, default=RUNTIME) args = parser.parse_args() root = args.root if args.phase in ("cache", "fit", "score-development", "score-protected"): record_preflight(evidence_root=root, phase=args.phase, storage_path=args.cache_root.parent) if args.phase == "freeze": result = freeze_authorities( config_path=CONFIG, split_manifest=SPLIT, reference_models=REFERENCE_ROOT / "models.json", reference_development=REFERENCE_ROOT / "direct-cubeful.json", canonical_manifest=CANONICAL / "manifest.json", output_path=root / "frozen-authorities.json", ) elif args.phase == "cache": result = build_training_cache( shallow_root=SHALLOW, split_manifest=SPLIT, cache_root=args.cache_root, output_path=root / "training-cache.json", ) elif args.phase == "fit": result = fit_models(cache_root=args.cache_root, output_path=root / "models.json") elif args.phase == "score-development": result = score_development( shallow_root=SHALLOW, split_manifest=SPLIT, reference_models=REFERENCE_ROOT / "models.json", reference_development=REFERENCE_ROOT / "direct-cubeful.json", models_path=root / "models.json", output_path=root / "development.json", contributions_path=root / "contribution-evidence.json", ) elif args.phase == "freeze-winner": result = freeze_winner( config_path=CONFIG, models_path=root / "models.json", development_path=root / "development.json", output_path=root / "winner-freeze.json", access_log_path=root / "protected-access-log.json", ) elif args.phase == "score-protected": result = score_protected( canonical_package=CANONICAL, models_path=root / "models.json", winner_path=root / "winner-freeze.json", access_log_path=root / "protected-access-log.json", output_path=root / "protected-final.json", ) elif args.phase == "summarize": result = build_summary( evidence_root=root, result_path=Path("results/improve-additive-native-cubeful-modeling-v1.json"), ) elif args.phase == "manifest": result = build_manifest(root) else: result = verify_package(root) print(json.dumps(result, indent=2, sort_keys=True)) return 0 if result.get("status", "PASS").startswith("PASS") else 1 if __name__ == "__main__": raise SystemExit(main()) from __future__ import annotations import json from pathlib import Path import numpy as np from backgammon_explainer.native_cubeful_experiment import ( ALPHA, DEVELOPMENT_FOLD_SEED, EPOCHS, FULL_WIDTH, MODEL_IDS, P3_WIDTH, AdditiveTransform, NativeCubefulModel, _bin_masks, _fold, ) CONFIG = Path("config/modeling/improve-additive-native-cubeful-modeling-v1.json") def test_frozen_config_and_absolute_boundaries() -> None: config = json.loads(CONFIG.read_text()) assert config["status"] == "FROZEN_BEFORE_DEVELOPMENT_SCORING" assert config["comparison"]["regularization_grid"] == [ALPHA] assert config["comparison"]["optimizer"]["epochs"] == EPOCHS assert config["data_authority"]["train"]["decisions"] == 1_000_002 assert config["data_authority"]["development"]["decisions"] == 100_015 assert config["data_authority"]["protected"]["access_before_frozen_winner"] == 0 assert config["activity_boundary"]["sage_gnu_campaign_training_rows"] == 0 assert config["activity_boundary"]["production_promotion"] is False assert config["calculated_cubeful_authority"] == "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" def test_frozen_model_shapes_and_exact_contributions() -> None: rng = np.random.default_rng(20260823) for model_id, width in zip(MODEL_IDS, (P3_WIDTH, FULL_WIDTH)): mean = rng.normal(size=width) scale = rng.uniform(0.5, 2.0, size=width) knots = np.sort(rng.normal(size=(width, 3)), axis=1) transform = AdditiveTransform(tuple(f"f{i}" for i in range(width)), mean, scale, knots) coefficients = rng.normal(size=width * 4) model = NativeCubefulModel(model_id, transform, coefficients, 0.25, {}) x = rng.normal(size=(3, FULL_WIDTH)) prediction = model.predict(x) contributions = model.grouped_contributions(x) assert contributions.shape == (3, width) assert np.allclose(prediction, contributions.sum(axis=1) + model.intercept, atol=1e-12) def test_development_fold_is_deterministic_and_bounded() -> None: first = [_fold(f"game-{index}") for index in range(100)] second = [_fold(f"game-{index}") for index in range(100)] assert DEVELOPMENT_FOLD_SEED == 20260823 assert first == second assert set(first) == {0, 1, 2, 3} def test_predeclared_segments_are_factual_and_exhaustive() -> None: context = np.zeros((4, 15)) context[:, 6] = [1, 2, 8, 16] context[:, 8] = [1, 0, 0, 1] context[:, 9] = [0, 1, 0, 0] context[:, 10] = [0, 0, 1, 0] context[:, 1] = [0, 3, 7, 15] context[:, 0] = [1, 0, 0, 0] context[:, 2] = [0, 0, 4, 2] context[:, 3] = [0, 0, 2, 5] context[:, 12] = [0, 1, 0, 1] truth = np.asarray([0.1, -0.3, 0.7, -1.2]) classes = np.asarray(["bar", "bearoff", "contact", "race"]) masks = _bin_masks(context, classes, truth) for family in ("target_magnitude", "cube_ownership", "score", "match_length", "crawford", "position_class"): selected = [mask for key, mask in masks.items() if key.startswith(family + "/")] assert np.asarray(selected, dtype=int).sum(axis=0).tolist() == [1, 1, 1, 1] def test_no_calculated_cubeful_or_competing_product_authority() -> None: source = Path("src/backgammon_explainer/native_cubeful_experiment.py").read_text() assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in source assert "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" in source assert '"production": "UNCHANGED"' in source { "version": "improve-additive-native-cubeful-modeling-v1", "status": "FROZEN_BEFORE_DEVELOPMENT_SCORING", "starting_implementation_head": "58522bb078ecda273a11476c60f1875a2255b285", "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", "accepted_integration_package_identity": "f40ba9417896383a94e48012843eb0f45e177430e746cd01080f8243c5751424", "calculated_cubeful_authority": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", "data_authority": { "source": "accepted GNU 0-ply modeling rows; no new engine work", "source_manifest_sha256": "a756567c6c6e316f0bf45527e516127af4317ab8e3b9ea5aa0156ca2dae1c14d", "split_manifest_identity_sha256": "eb0d571182b529588861731d29a49c26880d93f0170646b16e8974d81576ed0f", "split_manifest_sha256": "1d125f02d9c5c7340528e134ee6e2815d3ebe612e9df009ae236b726a92019d6", "train": { "checkpoint": "1000000", "complete_games": 32228, "decisions": 1000002, "candidates": 20981224, "membership_sha256": "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" }, "development": { "complete_games": 3178, "decisions": 100015, "candidates": 2094039, "membership_sha256": "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" }, "protected": { "authority": "existing actual-4ply non-adaptive final evaluation", "decisions": 2136, "candidates": 6963, "split_assignment_sha256": "7dfb06ff31d6623bca3d3832a1e12b52e12bd1c991fa0db81c357d892b3302a8", "canonical_manifest_sha256": "effa2a8bc273be03222d8193c090f96ef3415224af228f0e45678b8e7ec498a7", "access_before_frozen_winner": 0 }, "excluded_decision_membership_sha256": "c114765ae3e48330d928d665dbfdff4e902e90f2b44e0e1577006b35a161b6b9" }, "target": { "id": "native_cubeful_equity_static_next_player", "source_field": "native_equity", "transform": "-native_equity", "source_semantics": "GNU native checker-candidate Cubeful equity is maximized by the checker-move player", "modeled_perspective": "normalized static post-move next player on roll", "evaluation_mode": "Cubeful" }, "features": { "position_registry": "explainer-position-value-p3-v1-30ede35745bbbc64", "position_registry_sha256": "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287", "position_feature_count": 351, "context_feature_count": 15, "direct_registry_sha256": "b664347c564c0fecd53448554119941f94bb161075a61bf6cb957145949e604e", "context_order_sha256": "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914", "context_fields": [ "cubeful_is_money", "cubeful_match_length", "cubeful_player_score", "cubeful_opponent_score", "cubeful_player_away", "cubeful_opponent_away", "cubeful_cube_value", "cubeful_cube_log2", "cubeful_cube_centered", "cubeful_cube_owned_by_player", "cubeful_cube_owned_by_opponent", "cubeful_cube_owner_relative_code", "cubeful_crawford", "cubeful_jacoby", "cubeful_cube_offer_pending" ] }, "comparison": { "baseline": { "id": "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1", "model_identity_sha256": "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5", "models_artifact_sha256": "b27779b1947bcae91bddb067bd03d2267016140839f5b2bf9ce793ce6797c493", "accepted_development_evidence_sha256": "29d09310c078788685d653fb4b16e8fef84959d5ca1f6c002baad6f4eb532091", "feature_count": 366, "alpha": 10.0, "training_checkpoint": "full" }, "candidates": [ { "id": "native-cubeful-p3-additive-ridge-v1", "features": "P3 position only", "feature_count": 351 }, { "id": "native-cubeful-p3-context-additive-ridge-v1", "features": "P3 plus the complete accepted 15-field factual context block", "feature_count": 366 } ], "additive_basis": { "per_feature_terms": ["standardized_linear", "hinge_q25", "hinge_q50", "hinge_q75"], "hinge_quantiles": [0.25, 0.5, 0.75], "knots": "TRAIN-only", "interactions": 0, "intercept": true }, "regularization_grid": [100.0], "regularization_disposition": "singleton reuse of the accepted successful ADDEQ alpha; no outcome search", "optimizer": { "algorithm": "deterministic streaming Adam on the Ridge objective", "epochs": 12, "batch_size": 32768, "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, "beta1": 0.9, "beta2": 0.999, "training_arithmetic": "float32 basis/optimizer; float64 retained inference and reconstruction", "seed": 20260823 } }, "development_grouped_folds": { "version": "native-cubeful-development-complete-game-fold-4-v1", "seed": 20260823, "folds": 4, "assignment": "SHA256([version,seed,complete_game_id]) modulo four" }, "predeclared_bins": { "target_magnitude": ["abs<0.25", "0.25<=abs<0.5", "0.5<=abs<1", "abs>=1"], "cube_ownership": ["centered", "modeled_player", "opponent"], "cube_value": ["1", "2", "4", "8", "16+"], "score": ["money", "0-0", "tied_nonzero", "modeled_player_leads", "modeled_player_trails"], "match_length": ["money", "1-3", "4-7", "8-11", "12+"], "crawford": ["false", "true"], "position_class": ["bar", "bearoff", "contact", "race"] }, "decision_rule": { "minimum_absolute_rmse_improvement": 0.002, "minimum_absolute_mae_improvement": 0.002, "maximum_absolute_bias_regression": 0.005, "stable_fold_requirement": "improve both RMSE and MAE in at least 3 of 4 folds and regress neither by more than 0.002 in any fold", "winner_tie_break": ["lower development RMSE", "lower development MAE", "position-only before context"], "success_label": "MATERIAL_NATIVE_CUBEFUL_IMPROVEMENT", "protected_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" }, "software": { "python": "3.11.2", "numpy": "2.4.6", "scipy": "1.17.1", "scikit_learn": "1.9.0", "pyarrow": "25.0.1", "duckdb": "1.5.5" }, "host_policy": { "host": "high-fructose-corn-syrup", "nice": 19, "workers": 1, "blas_threads": 1, "fresh_preflight_before_each_substantial_phase": true, "protected_process_intervention": "FORBIDDEN" }, "activity_boundary": { "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, "new_generic_0ply": 0, "sage_gnu_campaign_training_rows": 0, "production_promotion": false, "analyzer_mutation": false, "canonical_mutation": false, "corpus_mutation": false } } exec /bin/bash -lc "sed -n '1,460p' src/backgammon_explainer/native_cubeful_experiment.py sed -n '461,920p' src/backgammon_explainer/native_cubeful_experiment.py" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 0ms: """Frozen existing-data experiment for additive GNU native Cubeful equity. The module deliberately exposes phase boundaries. In particular, development scoring cannot run before fitted models are durable and protected scoring cannot run before a development winner is frozen and logged. """ from __future__ import annotations import hashlib import json import math import os import platform import resource import socket import time from dataclasses import dataclass from datetime import datetime, timezone from pathlib import Path from typing import Any, Iterable, Mapping, Sequence import numpy as np import pyarrow.parquet as pq from .canonical_analysis import sha256_file, stable_json from .constrained_additive_position_model import HINGE_QUANTILES, _adam_step from .position_value_experiment import ( RegressionMetrics, SOURCE_COLUMNS, _candidate_files, _membership, _partition_key, load_frozen_deep_rows, load_models, position_classes, ) from .position_value_modeling import ( CUBEFUL_CONTEXT_REGISTRY, P3_REGISTRY, cubeful_context_matrix, position_feature_matrix, ) VERSION = "improve-additive-native-cubeful-modeling-v1" TRAIN_CHECKPOINT = "1000000" TRAIN_BUCKET_MAXIMUM = 4 EXPECTED_TRAIN_ROWS = 20_981_224 EXPECTED_DEVELOPMENT_ROWS = 2_094_039 EXPECTED_DEVELOPMENT_DECISIONS = 100_015 ALPHA = 100.0 EPOCHS = 12 BATCH_SIZE = 32_768 DEVELOPMENT_FOLD_VERSION = "native-cubeful-development-complete-game-fold-4-v1" DEVELOPMENT_FOLD_SEED = 20260823 MODEL_IDS = ( "native-cubeful-p3-additive-ridge-v1", "native-cubeful-p3-context-additive-ridge-v1", ) TARGET = "native_cubeful_equity_static_next_player" P3_WIDTH = len(P3_REGISTRY) FULL_WIDTH = P3_WIDTH + len(CUBEFUL_CONTEXT_REGISTRY) def _sha(value: Any) -> str: return hashlib.sha256(stable_json(value).encode()).hexdigest() def _write(path: Path, value: Any) -> None: path.parent.mkdir(parents=True, exist_ok=True) path.write_text(stable_json(value, pretty=True), encoding="utf-8") def _utc_now() -> str: return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") def _meminfo() -> dict[str, int]: result: dict[str, int] = {} for line in Path("/proc/meminfo").read_text().splitlines(): key, raw = line.split(":", 1) result[key] = int(raw.strip().split()[0]) return result def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: """Append the required fresh, read-only shared-host capacity observation.""" memory = _meminfo() stat = os.statvfs(storage_path) payload = { "recorded_at_utc": _utc_now(), "phase": phase, "host": socket.gethostname(), "load_average": list(os.getloadavg()), "logical_host_cpus": os.cpu_count(), "process_affinity_cpus": len(os.sched_getaffinity(0)), "mem_total_kib": memory["MemTotal"], "mem_available_kib": memory["MemAvailable"], "swap_total_kib": memory["SwapTotal"], "swap_free_kib": memory["SwapFree"], "storage_path": str(storage_path.resolve()), "storage_free_bytes": stat.f_bavail * stat.f_frsize, "storage_free_inodes": stat.f_favail, "nice": os.getpriority(os.PRIO_PROCESS, 0), "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", "protected_process_actions": [], "disposition": "PASS_SUBSTANTIAL_HEADROOM", } evidence_root.mkdir(parents=True, exist_ok=True) with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: stream.write(stable_json(payload) + "\n") return payload def freeze_authorities( *, config_path: Path, split_manifest: Path, reference_models: Path, reference_development: Path, canonical_manifest: Path, output_path: Path, ) -> dict[str, Any]: config = json.loads(config_path.read_text()) split = json.loads(split_manifest.read_text()) reference = json.loads(reference_models.read_text()) direct = json.loads(reference_development.read_text()) if config["status"] != "FROZEN_BEFORE_DEVELOPMENT_SCORING": raise RuntimeError("experiment configuration is not frozen") if sha256_file(split_manifest) != config["data_authority"]["split_manifest_sha256"]: raise RuntimeError("split authority hash differs") if split["manifest_identity_sha256"] != config["data_authority"]["split_manifest_identity_sha256"]: raise RuntimeError("split authority identity differs") if sha256_file(reference_models) != config["comparison"]["baseline"]["models_artifact_sha256"]: raise RuntimeError("baseline models hash differs") if sha256_file(reference_development) != config["comparison"]["baseline"]["accepted_development_evidence_sha256"]: raise RuntimeError("baseline development evidence hash differs") if sha256_file(canonical_manifest) != config["data_authority"]["protected"]["canonical_manifest_sha256"]: raise RuntimeError("protected manifest hash differs") model = [item for item in reference["models"] if item.get("model_identity_sha256") == config["comparison"]["baseline"]["model_identity_sha256"]] if len(model) != 1 or direct["model_identity_sha256"] != model[0]["model_identity_sha256"]: raise RuntimeError("unique accepted direct baseline was not found") checkpoint = split["selection"]["checkpoints"][TRAIN_CHECKPOINT] holdout = split["selection"]["holdout"] for observed, expected in ( (checkpoint["candidates"], EXPECTED_TRAIN_ROWS), (holdout["candidates"], EXPECTED_DEVELOPMENT_ROWS), (holdout["decisions"], EXPECTED_DEVELOPMENT_DECISIONS), ): if int(observed) != expected: raise RuntimeError("frozen population count differs") payload = { "version": VERSION + "-frozen-authorities-v1", "status": "PASS_FROZEN_BEFORE_DEVELOPMENT_SCORING", "config_path": str(config_path), "config_sha256": sha256_file(config_path), "starting_implementation_head": config["starting_implementation_head"], "source_and_partition_authority": config["data_authority"], "target_authority": config["target"], "feature_authority": config["features"], "model_comparison": config["comparison"], "development_grouped_folds": config["development_grouped_folds"], "predeclared_bins": config["predeclared_bins"], "decision_rule": config["decision_rule"], "software": config["software"], "host_policy": config["host_policy"], "activity_boundary": config["activity_boundary"], "accepted_baseline": { "model_identity_sha256": model[0]["model_identity_sha256"], "accepted_development_identity_sha256": direct["identity_sha256"], "accepted_metrics": direct["metrics"], }, "protected_accesses": [], } payload["identity_sha256"] = _sha(payload) _write(output_path, payload) return payload @dataclass class CachePart: x_t: np.ndarray y: np.ndarray indexes: np.ndarray path: Path def _cache_one(path: Path, games: Mapping[str, int], excluded: set[str], output: Path) -> dict[str, Any]: xs: list[np.ndarray] = [] ys: list[np.ndarray] = [] buckets: list[np.ndarray] = [] for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): data = batch.to_pydict() bucket = np.fromiter((games.get(str(value), -1) for value in data["game_key"]), dtype=np.int8) allowed = (bucket >= 0) & (bucket <= TRAIN_BUCKET_MAXIMUM) & np.fromiter( (str(value) not in excluded for value in data["decision_id"]), dtype=bool, ) chosen = np.flatnonzero(allowed) if not len(chosen): continue positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] match_ids = [str(data["source_match_id"][index]) for index in chosen] xs.append(np.column_stack(( position_feature_matrix(positions), cubeful_context_matrix(match_ids), )).astype(np.float32)) ys.append((-np.asarray(data["native_equity"], dtype=float)[chosen]).astype(np.float32)) buckets.append(bucket[chosen]) x = np.concatenate(xs) if xs else np.empty((0, FULL_WIDTH), dtype=np.float32) y = np.concatenate(ys) if ys else np.empty(0, dtype=np.float32) membership = np.concatenate(buckets) if buckets else np.empty(0, dtype=np.int8) output.mkdir(parents=True, exist_ok=True) np.save(output / "x_t.npy", x.T) np.save(output / "y.npy", y) np.save(output / "bucket.npy", membership) return { "source": str(path), "cache": str(output), "rows": len(y), "x_stream_sha256": hashlib.sha256(x.tobytes()).hexdigest(), "y_stream_sha256": hashlib.sha256(y.tobytes()).hexdigest(), "bucket_stream_sha256": hashlib.sha256(membership.tobytes()).hexdigest(), } def build_training_cache( *, shallow_root: Path, split_manifest: Path, cache_root: Path, output_path: Path, ) -> dict[str, Any]: started = time.time() split = json.loads(split_manifest.read_text()) train, _, excluded = _membership(split) records = [] for index, path in enumerate(_candidate_files(shallow_root)): records.append(_cache_one( path, train.get(_partition_key(path), {}), excluded, cache_root / f"partition-{index:03d}", )) rows = sum(item["rows"] for item in records) if rows != EXPECTED_TRAIN_ROWS: raise RuntimeError(f"cached training population differs: {rows}") payload = { "version": VERSION + "-training-cache-v1", "status": "PASS", "source_split_identity_sha256": split["manifest_identity_sha256"], "checkpoint": TRAIN_CHECKPOINT, "candidate_rows": rows, "partition_count": len(records), "parts": records, "cache_root": str(cache_root.resolve()), "elapsed_seconds": time.time() - started, "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, "activity_boundary": {"new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0}, } payload["deterministic_identity_sha256"] = _sha({ key: payload[key] for key in ("version", "status", "source_split_identity_sha256", "checkpoint", "candidate_rows", "partition_count", "parts") }) _write(output_path, payload) return payload def load_cache(cache_root: Path) -> list[CachePart]: parts = [] for path in sorted(cache_root.glob("partition-*")): x_t = np.load(path / "x_t.npy", mmap_mode="r") y = np.load(path / "y.npy", mmap_mode="r") bucket = np.load(path / "bucket.npy", mmap_mode="r") indexes = np.flatnonzero(bucket <= TRAIN_BUCKET_MAXIMUM) if x_t.shape != (FULL_WIDTH, len(y)) or len(bucket) != len(y): raise RuntimeError(f"invalid cache part {path}") parts.append(CachePart(x_t, y, indexes, path)) if len(parts) != 82 or sum(len(part.indexes) for part in parts) != EXPECTED_TRAIN_ROWS: raise RuntimeError("training cache population differs") return parts def _iter_cache(parts: Sequence[CachePart], *, batch_size: int = BATCH_SIZE) -> Iterable[tuple[np.ndarray, np.ndarray]]: for part in parts: for start in range(0, len(part.indexes), batch_size): chosen = part.indexes[start:start + batch_size] yield np.asarray(part.x_t[:, chosen].T, dtype=np.float32), np.asarray(part.y[chosen], dtype=np.float32) @dataclass class AdditiveTransform: feature_ids: tuple[str, ...] mean: np.ndarray scale: np.ndarray knots: np.ndarray @property def width(self) -> int: return len(self.feature_ids) @property def basis_width(self) -> int: return self.width * 4 def basis_float32(self, x: np.ndarray) -> np.ndarray: z = (np.asarray(x[:, :self.width], dtype=np.float32) - self.mean.astype(np.float32)) / self.scale.astype(np.float32) output = np.empty((len(z), self.width, 4), dtype=np.float32) output[:, :, 0] = z output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) return output.reshape(len(z), -1) def basis(self, x: np.ndarray) -> np.ndarray: z = (np.asarray(x[:, :self.width], dtype=float) - self.mean) / self.scale output = np.empty((len(z), self.width, 4), dtype=float) output[:, :, 0] = z output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) return output.reshape(len(z), -1) def descriptor(self) -> dict[str, Any]: return { "feature_ids": list(self.feature_ids), "standard_scaler_mean": self.mean.tolist(), "standard_scaler_scale": self.scale.tolist(), "hinge_quantiles": list(HINGE_QUANTILES), "hinge_knots_standardized": self.knots.tolist(), "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", "interactions": 0, } def fit_transform(parts: Sequence[CachePart]) -> AdditiveTransform: total = 0 sums = np.zeros(FULL_WIDTH) squares = np.zeros(FULL_WIDTH) for x, _ in _iter_cache(parts, batch_size=8192): total += len(x) sums += x.sum(axis=0, dtype=float) squares += np.square(x, dtype=float).sum(axis=0) mean = sums / total scale = np.sqrt(np.maximum(0.0, squares / total - mean * mean)) scale[scale == 0] = 1.0 raw_knots = np.empty((FULL_WIDTH, 3)) for feature in range(FULL_WIDTH): columns = [np.asarray(part.x_t[feature, part.indexes]) for part in parts] raw_knots[feature] = np.quantile(np.concatenate(columns), HINGE_QUANTILES) knots = (raw_knots - mean[:, None]) / scale[:, None] feature_ids = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) return AdditiveTransform(feature_ids, mean, scale, knots) @dataclass class NativeCubefulModel: model_id: str transform: AdditiveTransform coefficients: np.ndarray intercept: float optimizer: dict[str, Any] def predict(self, x: np.ndarray) -> np.ndarray: return self.transform.basis(x) @ self.coefficients + self.intercept def grouped_contributions(self, x: np.ndarray) -> np.ndarray: basis = self.transform.basis(x) return (basis * self.coefficients).reshape(len(x), self.transform.width, 4).sum(axis=2) def descriptor(self) -> dict[str, Any]: payload = { "model_id": self.model_id, "target": TARGET, "alpha": ALPHA, "training_checkpoint": TRAIN_CHECKPOINT, "transform": self.transform.descriptor(), "coefficients": self.coefficients.tolist(), "intercept": self.intercept, "optimizer": self.optimizer, "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", } payload["model_identity_sha256"] = _sha(payload) return payload def _slice_transform(full: AdditiveTransform, width: int) -> AdditiveTransform: return AdditiveTransform( full.feature_ids[:width], full.mean[:width], full.scale[:width], full.knots[:width], ) def fit_models(*, cache_root: Path, output_path: Path) -> dict[str, Any]: parts = load_cache(cache_root) started = time.time() full_transform = fit_transform(parts) transforms = (_slice_transform(full_transform, P3_WIDTH), full_transform) rows = sum(len(part.indexes) for part in parts) target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows parameters = [np.zeros(transform.basis_width + 1, dtype=np.float32) for transform in transforms] for parameter in parameters: parameter[-1] = target_mean first = [np.zeros_like(parameter) for parameter in parameters] second = [np.zeros_like(parameter) for parameter in parameters] step = 0 epoch_records = [] for epoch in range(EPOCHS): sums_squared = np.zeros(2) seen = 0 maximum_updates = np.zeros(2) for x, y in _iter_cache(parts): full_basis = full_transform.basis_float32(x) bases = (full_basis[:, :P3_WIDTH * 4], full_basis) for index, (basis, parameter) in enumerate(zip(bases, parameters)): residual = basis @ parameter[:-1] + parameter[-1] - y sums_squared[index] += float(np.square(residual).sum()) gradient = np.append(residual @ basis / len(x), residual.mean()).astype(np.float32) gradient[:-1] += (ALPHA / rows) * parameter[:-1] batch_lr = 0.015 * (0.75 ** epoch) * min(1.0, len(x) / BATCH_SIZE) maximum_updates[index] = max( maximum_updates[index], _adam_step(parameter, gradient, first[index], second[index], step + 1, batch_lr), ) step += 1 seen += len(x) epoch_records.append({ "epoch": epoch + 1, "online_rmse": { MODEL_IDS[index]: math.sqrt(sums_squared[index] / seen) for index in range(2) }, "maximum_absolute_parameter_update": { MODEL_IDS[index]: float(maximum_updates[index]) for index in range(2) }, }) common = { "algorithm": "deterministic streaming Adam on Ridge objective", "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, "beta1": 0.9, "beta2": 0.999, "seed": 20260823, "training_rows": rows, "updates": step, "epoch_records": epoch_records, "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", } models = [ NativeCubefulModel(MODEL_IDS[index], transform, parameter[:-1].astype(float), float(parameter[-1]), {**common, "model_index": index}) for index, (transform, parameter) in enumerate(zip(transforms, parameters)) ] payload = { "version": VERSION + "-models-v1", "status": "PASS_FROZEN_BEFORE_DEVELOPMENT", "models": [model.descriptor() for model in models], "elapsed_seconds": time.time() - started, "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, "development_accesses": 0, "protected_accesses": 0, } payload["deterministic_identity_sha256"] = _sha({ "version": payload["version"], "status": payload["status"], "models": payload["models"], "development_accesses": 0, "protected_accesses": 0, }) _write(output_path, payload) return payload def load_native_models(path: Path) -> list[NativeCubefulModel]: payload = json.loads(path.read_text()) result = [] for item in payload["models"]: transform = item["transform"] model = NativeCubefulModel( item["model_id"], AdditiveTransform( tuple(transform["feature_ids"]), np.asarray(transform["standard_scaler_mean"]), np.asarray(transform["standard_scaler_scale"]), np.asarray(transform["hinge_knots_standardized"]), ), np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), ) if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: raise RuntimeError("model identity differs") result.append(model) return result def _fold(game_id: str) -> int: digest = _sha([DEVELOPMENT_FOLD_VERSION, DEVELOPMENT_FOLD_SEED, game_id]) return int(digest[:16], 16) % 4 def _bin_masks(context: np.ndarray, classes: np.ndarray, truth: np.ndarray) -> dict[str, np.ndarray]: absolute = np.abs(truth) result = { "target_magnitude/abs<0.25": absolute < 0.25, "target_magnitude/0.25<=abs<0.5": (absolute >= 0.25) & (absolute < 0.5), "target_magnitude/0.5<=abs<1": (absolute >= 0.5) & (absolute < 1.0), "target_magnitude/abs>=1": absolute >= 1.0, "cube_ownership/centered": context[:, 8] == 1, "cube_ownership/modeled_player": context[:, 9] == 1, "cube_ownership/opponent": context[:, 10] == 1, "crawford/false": context[:, 12] == 0, "crawford/true": context[:, 12] == 1, } cube = context[:, 6] for value in (1, 2, 4, 8): result[f"cube_value/{value}"] = cube == value result["cube_value/16+"] = cube >= 16 money, match_length = context[:, 0] == 1, context[:, 1] player_score, opponent_score = context[:, 2], context[:, 3] result.update({ "score/money": money, "score/0-0": (~money) & (player_score == 0) & (opponent_score == 0), "score/tied_nonzero": (~money) & (player_score == opponent_score) & (player_score > 0), "score/modeled_player_leads": (~money) & (player_score > opponent_score), "score/modeled_player_trails": (~money) & (player_score < opponent_score), "match_length/money": money, "match_length/1-3": (~money) & (match_length <= 3), "match_length/4-7": (match_length >= 4) & (match_length <= 7), "match_length/8-11": (match_length >= 8) & (match_length <= 11), "match_length/12+": match_length >= 12, }) for label in ("bar", "bearoff", "contact", "race"): result[f"position_class/{label}"] = classes == label return result class DetailedMetrics: def __init__(self) -> None: self.global_metric = RegressionMetrics() self.folds = {index: RegressionMetrics() for index in range(4)} self.segments: dict[str, RegressionMetrics] = {} def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray, masks: Mapping[str, np.ndarray]) -> None: self.global_metric.add(prediction, truth) for fold in range(4): selected = folds == fold if np.any(selected): self.folds[fold].add(prediction[selected], truth[selected]) for label, selected in masks.items(): if np.any(selected): self.segments.setdefault(label, RegressionMetrics()).add(prediction[selected], truth[selected]) def result(self) -> dict[str, Any]: return { "global": self.global_metric.result(), "development_complete_game_folds": { str(index): metric.result() for index, metric in self.folds.items() }, "segments": {label: metric.result() for label, metric in sorted(self.segments.items())}, } def _accepted_baseline(reference_models: Path): matches = [model for model in load_models(reference_models) if model.feature_set == "P3+CUBEFUL_CONTEXT" and model.checkpoint == "full"] if len(matches) != 1 or matches[0].alpha != 10.0: raise RuntimeError("unique accepted direct Cubeful baseline missing") return matches[0] def score_development( *, shallow_root: Path, split_manifest: Path, reference_models: Path, reference_development: Path, models_path: Path, output_path: Path, contributions_path: Path, ) -> dict[str, Any]: if not models_path.exists(): raise RuntimeError("models must be durable before DEVELOPMENT access") split = json.loads(split_manifest.read_text()) _, holdout, _ = _membership(split) campaign = str(split["source_authority"]["campaign"]) native_models = load_native_models(models_path) baseline = _accepted_baseline(reference_models) metrics = {"accepted_baseline": DetailedMetrics(), **{ model.model_id: DetailedMetrics() for model in native_models }} samples: tuple[np.ndarray, list[str], np.ndarray] | None = None rows = 0 decisions: set[str] = set() started = time.time() for path in _candidate_files(shallow_root): games = holdout.get(_partition_key(path), set()) if not games: continue host, worker = _partition_key(path) for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): data = batch.to_pydict() chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) if not len(chosen): continue positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] match_ids = [str(data["source_match_id"][index]) for index in chosen] context = cubeful_context_matrix(match_ids) x = np.column_stack((position_feature_matrix(positions), context)) truth = -np.asarray(data["native_equity"], dtype=float)[chosen] classes = position_classes(positions) folds = np.fromiter(( _fold("\0".join((campaign, host, worker, str(data["game_key"][index])))) for index in chosen ), dtype=np.int8) masks = _bin_masks(context, classes, truth) metrics["accepted_baseline"].add(baseline.predict(x)[:, 0], truth, folds, masks) for model in native_models: metrics[model.model_id].add(model.predict(x), truth, folds, masks) if samples is None: take = chosen[:2] samples = (x[:len(take)].copy(), [str(data["decision_id"][index]) for index in take], truth[:len(take)].copy()) rows += len(chosen) decisions.update(str(data["decision_id"][index]) for index in chosen) if rows != EXPECTED_DEVELOPMENT_ROWS or len(decisions) != EXPECTED_DEVELOPMENT_DECISIONS: raise RuntimeError("DEVELOPMENT population differs") observed = metrics["accepted_baseline"].result() accepted = json.loads(reference_development.read_text())["metrics"] reproduction = {key: abs(observed["global"][key] - accepted[key]) for key in ("rmse", "mae", "bias", "r2", "correlation")} if max(reproduction.values()) > 1e-12: raise RuntimeError(f"accepted baseline did not reproduce: {reproduction}") if samples is None: raise RuntimeError("no DEVELOPMENT explanation sample") sample_x, decision_ids, sample_truth = samples explanation_models = {} maximum_error = 0.0 for model in native_models: prediction = model.predict(sample_x) contributions = model.grouped_contributions(sample_x) reconstructed = contributions.sum(axis=1) + model.intercept error = float(np.max(np.abs(prediction - reconstructed))) maximum_error = max(maximum_error, error) explanation_models[model.model_id] = { "model_identity_sha256": model.descriptor()["model_identity_sha256"], "decision_ids": decision_ids, "truth": sample_truth.tolist(), "prediction": prediction.tolist(), "intercept": model.intercept, "feature_contributions": [ {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} for index, feature_id in enumerate(model.transform.feature_ids) ], "position_reconstruction_maximum_absolute_error": error, "a_minus_b_reconstruction_absolute_error": abs( float((prediction[0] - prediction[1]) - (contributions[0] - contributions[1]).sum()) ), } if maximum_error > 1e-10: raise RuntimeError("additive contribution reconstruction failed") contribution_payload = { "version": VERSION + "-contribution-evidence-v1", "status": "PASS", "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", "models": explanation_models, "global_maximum_absolute_reconstruction_error": maximum_error, "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", } contribution_payload["identity_sha256"] = _sha(contribution_payload) _write(contributions_path, contribution_payload) payload = { "version": VERSION + "-development-v1", "status": "PASS", "access": "DEVELOPMENT", "candidates": rows, "decisions": len(decisions), "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, "models_identity_sha256": json.loads(models_path.read_text())["deterministic_identity_sha256"], "metrics": {name: metric.result() for name, metric in metrics.items()}, "accepted_baseline_reproduction_absolute_error": reproduction, "contribution_evidence_identity_sha256": contribution_payload["identity_sha256"], "elapsed_seconds": time.time() - started, "protected_accesses": 0, } payload["identity_sha256"] = _sha(payload) _write(output_path, payload) return payload def freeze_winner(*, config_path: Path, models_path: Path, development_path: Path, output_path: Path, access_log_path: Path) -> dict[str, Any]: config = json.loads(config_path.read_text()) development = json.loads(development_path.read_text()) models = json.loads(models_path.read_text()) baseline = development["metrics"]["accepted_baseline"] rule = config["decision_rule"] qualifying = [] audits = {} by_id = {item["model_id"]: item for item in models["models"]} for model_id in MODEL_IDS: candidate = development["metrics"][model_id] fold_improvements = 0 maximum_fold_rmse_regression = -math.inf maximum_fold_mae_regression = -math.inf for fold in range(4): base_fold = baseline["development_complete_game_folds"][str(fold)] cand_fold = candidate["development_complete_game_folds"][str(fold)] if cand_fold["rmse"] < base_fold["rmse"] and cand_fold["mae"] < base_fold["mae"]: fold_improvements += 1 maximum_fold_rmse_regression = max(maximum_fold_rmse_regression, cand_fold["rmse"] - base_fold["rmse"]) maximum_fold_mae_regression = max(maximum_fold_mae_regression, cand_fold["mae"] - base_fold["mae"]) rmse_gain = baseline["global"]["rmse"] - candidate["global"]["rmse"] mae_gain = baseline["global"]["mae"] - candidate["global"]["mae"] bias_regression = abs(candidate["global"]["bias"]) - abs(baseline["global"]["bias"]) passed = ( rmse_gain >= rule["minimum_absolute_rmse_improvement"] and mae_gain >= rule["minimum_absolute_mae_improvement"] and bias_regression <= rule["maximum_absolute_bias_regression"] and fold_improvements >= 3 and maximum_fold_rmse_regression <= 0.002 and maximum_fold_mae_regression <= 0.002 ) audits[model_id] = { "rmse_improvement": rmse_gain, "mae_improvement": mae_gain, "absolute_bias_regression": bias_regression, "folds_improving_both": fold_improvements, "maximum_fold_rmse_regression": maximum_fold_rmse_regression, "maximum_fold_mae_regression": maximum_fold_mae_regression, "passes_material_gate": passed, } if passed: qualifying.append(model_id) qualifying.sort(key=lambda model_id: ( development["metrics"][model_id]["global"]["rmse"], development["metrics"][model_id]["global"]["mae"], MODEL_IDS.index(model_id), )) winner = qualifying[0] if qualifying else None payload = { "version": VERSION + "-winner-freeze-v1", "status": "MATERIAL_NATIVE_CUBEFUL_IMPROVEMENT" if winner else "MATERIAL_NATIVE_CUBEFUL_IMPROVEMENT_NOT_ESTABLISHED", "development_identity_sha256": development["identity_sha256"], "candidate_audits": audits, "qualifying_candidates": qualifying, "winner_model_id": winner, "winner_model_identity_sha256": by_id[winner]["model_identity_sha256"] if winner else None, "winner_frozen_before_protected_access": True, "protected_accesses_at_freeze": 0, "production": "UNCHANGED", "accepted_product_architecture": config["accepted_product_architecture"], "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if winner else None, } payload["identity_sha256"] = _sha(payload) _write(output_path, payload) access_log = { "version": VERSION + "-protected-access-log-v1", "winner_freeze_identity_sha256": payload["identity_sha256"], "winner_model_identity_sha256": payload["winner_model_identity_sha256"], "accesses": [], } access_log["identity_sha256"] = _sha(access_log) _write(access_log_path, access_log) return payload def score_protected(*, canonical_package: Path, models_path: Path, winner_path: Path, access_log_path: Path, output_path: Path) -> dict[str, Any]: winner = json.loads(winner_path.read_text()) if not winner["winner_model_id"]: payload = { "version": VERSION + "-protected-v1", "status": "NOT_ACCESSED_NO_DEVELOPMENT_WINNER", "winner_freeze_identity_sha256": winner["identity_sha256"], "accesses": 0, } payload["identity_sha256"] = _sha(payload) _write(output_path, payload) return payload access = json.loads(access_log_path.read_text()) if access["accesses"]: raise RuntimeError("PROTECTED authority already accessed; refusing a second access") models = {model.model_id: model for model in load_native_models(models_path)} model = models[winner["winner_model_id"]] if model.descriptor()["model_identity_sha256"] != winner["winner_model_identity_sha256"]: raise RuntimeError("frozen winner identity differs") record = { "ordinal": 1, "started_at_utc": _utc_now(), "completed_at_utc": None, "purpose": "single final evaluation of the frozen DEVELOPMENT winner", "winner_model_id": model.model_id, "winner_model_identity_sha256": winner["winner_model_identity_sha256"], "canonical_manifest_sha256": sha256_file(canonical_package / "manifest.json"), "status": "STARTED_BEFORE_ROW_READ", } access["accesses"].append(record) access.pop("identity_sha256", None) access["identity_sha256"] = _sha(access) _write(access_log_path, access) rows = load_frozen_deep_rows(canonical_package) positions = [str(row["gnu_position_id"]) for row in rows] context = cubeful_context_matrix([str(row["gnu_match_id_native"]) for row in rows]) x = np.column_stack((position_feature_matrix(positions), context)) truth = -np.asarray([row["native_equity"] for row in rows], dtype=float) prediction = model.predict(x) metric = RegressionMetrics() metric.add(prediction, truth) segments = {} for label, mask in _bin_masks(context, position_classes(positions), truth).items(): if np.any(mask): item = RegressionMetrics(); item.add(prediction[mask], truth[mask]); segments[label] = item.result() payload = { "version": VERSION + "-protected-v1", "status": "PASS_SINGLE_FROZEN_WINNER_ACCESS", "winner_model_id": model.model_id, "winner_model_identity_sha256": winner["winner_model_identity_sha256"], "winner_freeze_identity_sha256": winner["identity_sha256"], "candidates": len(rows), "decisions": len({row["decision_id"] for row in rows}), "metrics": metric.result(), "segments": dict(sorted(segments.items())), "disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION", "production": "UNCHANGED", } payload["identity_sha256"] = _sha(payload) _write(output_path, payload) record["completed_at_utc"] = _utc_now() record["status"] = "COMPLETED" record["candidate_rows"] = len(rows) record["decision_rows"] = len({row["decision_id"] for row in rows}) record["result_identity_sha256"] = payload["identity_sha256"] access.pop("identity_sha256", None) access["identity_sha256"] = _sha(access) _write(access_log_path, access) return payload def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: authorities = json.loads((evidence_root / "frozen-authorities.json").read_text()) models = json.loads((evidence_root / "models.json").read_text()) development = json.loads((evidence_root / "development.json").read_text()) winner = json.loads((evidence_root / "winner-freeze.json").read_text()) protected = json.loads((evidence_root / "protected-final.json").read_text()) access = json.loads((evidence_root / "protected-access-log.json").read_text()) payload = { "version": VERSION + "-result-v1", "status": "PASS_COMPLETE", "starting_implementation_head": authorities["starting_implementation_head"], "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", "accepted_integration_package_identity": "f40ba9417896383a94e48012843eb0f45e177430e746cd01080f8243c5751424", "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", "partitions": {"train": authorities["source_and_partition_authority"]["train"], "development": authorities["source_and_partition_authority"]["development"], "protected": authorities["source_and_partition_authority"]["protected"]}, "target": authorities["target_authority"], "model_identities": {item["model_id"]: item["model_identity_sha256"] for item in models["models"]}, "development_metrics": {name: value["global"] for name, value in development["metrics"].items()}, "development_fold_metrics": {name: value["development_complete_game_folds"] for name, value in development["metrics"].items()}, "winner": winner, "protected_final": protected, "protected_access_log_identity_sha256": access["identity_sha256"], "protected_access_count": len(access["accesses"]), "production": "UNCHANGED", "analyzer": "UNCHANGED", "canonical": "UNCHANGED", "corpus": "UNCHANGED", "activity_boundary": authorities["activity_boundary"], "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", } payload["identity_sha256"] = _sha(payload) _write(evidence_root / "result-summary.json", payload) _write(result_path, payload) return payload def build_manifest(evidence_root: Path) -> dict[str, Any]: excluded = {"manifest.json", "SHA256SUMS", "self-verification.json", "preflight-log.jsonl"} files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) entries = [{"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} for path in files] payload = {"version": VERSION + "-manifest-v1", "files": entries} payload["package_identity_sha256"] = _sha(payload) _write(evidence_root / "manifest.json", payload) (evidence_root / "SHA256SUMS").write_text( "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", ) return payload def verify_package(evidence_root: Path) -> dict[str, Any]: manifest = json.loads((evidence_root / "manifest.json").read_text()) checks = [] for item in manifest["files"]: path = evidence_root / item["path"] checks.append({"name": "hash:" + item["path"], "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL"}) winner = json.loads((evidence_root / "winner-freeze.json").read_text()) access = json.loads((evidence_root / "protected-access-log.json").read_text()) protected = json.loads((evidence_root / "protected-final.json").read_text()) checks.extend(( {"name": "winner frozen before protected", "status": "PASS" if winner["winner_frozen_before_protected_access"] else "FAIL"}, {"name": "protected access count bounded", "status": "PASS" if len(access["accesses"]) <= 1 else "FAIL"}, {"name": "production unchanged", "status": "PASS" if winner["production"] == "UNCHANGED" else "FAIL"}, {"name": "calculated cubeful blocked", "status": "PASS" if json.loads((evidence_root / "result-summary.json").read_text())["calculated_cubeful"] == "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" else "FAIL"}, {"name": "protected disposition bounded", "status": "PASS" if protected.get("disposition") in (None, "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION") else "FAIL"}, )) payload = { "version": VERSION + "-self-verification-v1", "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", "package_identity_sha256": manifest["package_identity_sha256"], "checks": checks, "python": platform.python_version(), } _write(evidence_root / "self-verification.json", payload) return payload exec /bin/bash -lc "sed -n '220,560p' src/backgammon_explainer/native_cubeful_experiment.py" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 0ms: def build_training_cache( *, shallow_root: Path, split_manifest: Path, cache_root: Path, output_path: Path, ) -> dict[str, Any]: started = time.time() split = json.loads(split_manifest.read_text()) train, _, excluded = _membership(split) records = [] for index, path in enumerate(_candidate_files(shallow_root)): records.append(_cache_one( path, train.get(_partition_key(path), {}), excluded, cache_root / f"partition-{index:03d}", )) rows = sum(item["rows"] for item in records) if rows != EXPECTED_TRAIN_ROWS: raise RuntimeError(f"cached training population differs: {rows}") payload = { "version": VERSION + "-training-cache-v1", "status": "PASS", "source_split_identity_sha256": split["manifest_identity_sha256"], "checkpoint": TRAIN_CHECKPOINT, "candidate_rows": rows, "partition_count": len(records), "parts": records, "cache_root": str(cache_root.resolve()), "elapsed_seconds": time.time() - started, "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, "activity_boundary": {"new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0}, } payload["deterministic_identity_sha256"] = _sha({ key: payload[key] for key in ("version", "status", "source_split_identity_sha256", "checkpoint", "candidate_rows", "partition_count", "parts") }) _write(output_path, payload) return payload def load_cache(cache_root: Path) -> list[CachePart]: parts = [] for path in sorted(cache_root.glob("partition-*")): x_t = np.load(path / "x_t.npy", mmap_mode="r") y = np.load(path / "y.npy", mmap_mode="r") bucket = np.load(path / "bucket.npy", mmap_mode="r") indexes = np.flatnonzero(bucket <= TRAIN_BUCKET_MAXIMUM) if x_t.shape != (FULL_WIDTH, len(y)) or len(bucket) != len(y): raise RuntimeError(f"invalid cache part {path}") parts.append(CachePart(x_t, y, indexes, path)) if len(parts) != 82 or sum(len(part.indexes) for part in parts) != EXPECTED_TRAIN_ROWS: raise RuntimeError("training cache population differs") return parts def _iter_cache(parts: Sequence[CachePart], *, batch_size: int = BATCH_SIZE) -> Iterable[tuple[np.ndarray, np.ndarray]]: for part in parts: for start in range(0, len(part.indexes), batch_size): chosen = part.indexes[start:start + batch_size] yield np.asarray(part.x_t[:, chosen].T, dtype=np.float32), np.asarray(part.y[chosen], dtype=np.float32) @dataclass class AdditiveTransform: feature_ids: tuple[str, ...] mean: np.ndarray scale: np.ndarray knots: np.ndarray @property def width(self) -> int: return len(self.feature_ids) @property def basis_width(self) -> int: return self.width * 4 def basis_float32(self, x: np.ndarray) -> np.ndarray: z = (np.asarray(x[:, :self.width], dtype=np.float32) - self.mean.astype(np.float32)) / self.scale.astype(np.float32) output = np.empty((len(z), self.width, 4), dtype=np.float32) output[:, :, 0] = z output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) return output.reshape(len(z), -1) def basis(self, x: np.ndarray) -> np.ndarray: z = (np.asarray(x[:, :self.width], dtype=float) - self.mean) / self.scale output = np.empty((len(z), self.width, 4), dtype=float) output[:, :, 0] = z output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) return output.reshape(len(z), -1) def descriptor(self) -> dict[str, Any]: return { "feature_ids": list(self.feature_ids), "standard_scaler_mean": self.mean.tolist(), "standard_scaler_scale": self.scale.tolist(), "hinge_quantiles": list(HINGE_QUANTILES), "hinge_knots_standardized": self.knots.tolist(), "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", "interactions": 0, } def fit_transform(parts: Sequence[CachePart]) -> AdditiveTransform: total = 0 sums = np.zeros(FULL_WIDTH) squares = np.zeros(FULL_WIDTH) for x, _ in _iter_cache(parts, batch_size=8192): total += len(x) sums += x.sum(axis=0, dtype=float) squares += np.square(x, dtype=float).sum(axis=0) mean = sums / total scale = np.sqrt(np.maximum(0.0, squares / total - mean * mean)) scale[scale == 0] = 1.0 raw_knots = np.empty((FULL_WIDTH, 3)) for feature in range(FULL_WIDTH): columns = [np.asarray(part.x_t[feature, part.indexes]) for part in parts] raw_knots[feature] = np.quantile(np.concatenate(columns), HINGE_QUANTILES) knots = (raw_knots - mean[:, None]) / scale[:, None] feature_ids = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) return AdditiveTransform(feature_ids, mean, scale, knots) @dataclass class NativeCubefulModel: model_id: str transform: AdditiveTransform coefficients: np.ndarray intercept: float optimizer: dict[str, Any] def predict(self, x: np.ndarray) -> np.ndarray: return self.transform.basis(x) @ self.coefficients + self.intercept def grouped_contributions(self, x: np.ndarray) -> np.ndarray: basis = self.transform.basis(x) return (basis * self.coefficients).reshape(len(x), self.transform.width, 4).sum(axis=2) def descriptor(self) -> dict[str, Any]: payload = { "model_id": self.model_id, "target": TARGET, "alpha": ALPHA, "training_checkpoint": TRAIN_CHECKPOINT, "transform": self.transform.descriptor(), "coefficients": self.coefficients.tolist(), "intercept": self.intercept, "optimizer": self.optimizer, "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", } payload["model_identity_sha256"] = _sha(payload) return payload def _slice_transform(full: AdditiveTransform, width: int) -> AdditiveTransform: return AdditiveTransform( full.feature_ids[:width], full.mean[:width], full.scale[:width], full.knots[:width], ) def fit_models(*, cache_root: Path, output_path: Path) -> dict[str, Any]: parts = load_cache(cache_root) started = time.time() full_transform = fit_transform(parts) transforms = (_slice_transform(full_transform, P3_WIDTH), full_transform) rows = sum(len(part.indexes) for part in parts) target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows parameters = [np.zeros(transform.basis_width + 1, dtype=np.float32) for transform in transforms] for parameter in parameters: parameter[-1] = target_mean first = [np.zeros_like(parameter) for parameter in parameters] second = [np.zeros_like(parameter) for parameter in parameters] step = 0 epoch_records = [] for epoch in range(EPOCHS): sums_squared = np.zeros(2) seen = 0 maximum_updates = np.zeros(2) for x, y in _iter_cache(parts): full_basis = full_transform.basis_float32(x) bases = (full_basis[:, :P3_WIDTH * 4], full_basis) for index, (basis, parameter) in enumerate(zip(bases, parameters)): residual = basis @ parameter[:-1] + parameter[-1] - y sums_squared[index] += float(np.square(residual).sum()) gradient = np.append(residual @ basis / len(x), residual.mean()).astype(np.float32) gradient[:-1] += (ALPHA / rows) * parameter[:-1] batch_lr = 0.015 * (0.75 ** epoch) * min(1.0, len(x) / BATCH_SIZE) maximum_updates[index] = max( maximum_updates[index], _adam_step(parameter, gradient, first[index], second[index], step + 1, batch_lr), ) step += 1 seen += len(x) epoch_records.append({ "epoch": epoch + 1, "online_rmse": { MODEL_IDS[index]: math.sqrt(sums_squared[index] / seen) for index in range(2) }, "maximum_absolute_parameter_update": { MODEL_IDS[index]: float(maximum_updates[index]) for index in range(2) }, }) common = { "algorithm": "deterministic streaming Adam on Ridge objective", "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, "beta1": 0.9, "beta2": 0.999, "seed": 20260823, "training_rows": rows, "updates": step, "epoch_records": epoch_records, "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", } models = [ NativeCubefulModel(MODEL_IDS[index], transform, parameter[:-1].astype(float), float(parameter[-1]), {**common, "model_index": index}) for index, (transform, parameter) in enumerate(zip(transforms, parameters)) ] payload = { "version": VERSION + "-models-v1", "status": "PASS_FROZEN_BEFORE_DEVELOPMENT", "models": [model.descriptor() for model in models], "elapsed_seconds": time.time() - started, "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, "development_accesses": 0, "protected_accesses": 0, } payload["deterministic_identity_sha256"] = _sha({ "version": payload["version"], "status": payload["status"], "models": payload["models"], "development_accesses": 0, "protected_accesses": 0, }) _write(output_path, payload) return payload def load_native_models(path: Path) -> list[NativeCubefulModel]: payload = json.loads(path.read_text()) result = [] for item in payload["models"]: transform = item["transform"] model = NativeCubefulModel( item["model_id"], AdditiveTransform( tuple(transform["feature_ids"]), np.asarray(transform["standard_scaler_mean"]), np.asarray(transform["standard_scaler_scale"]), np.asarray(transform["hinge_knots_standardized"]), ), np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), ) if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: raise RuntimeError("model identity differs") result.append(model) return result def _fold(game_id: str) -> int: digest = _sha([DEVELOPMENT_FOLD_VERSION, DEVELOPMENT_FOLD_SEED, game_id]) return int(digest[:16], 16) % 4 def _bin_masks(context: np.ndarray, classes: np.ndarray, truth: np.ndarray) -> dict[str, np.ndarray]: absolute = np.abs(truth) result = { "target_magnitude/abs<0.25": absolute < 0.25, "target_magnitude/0.25<=abs<0.5": (absolute >= 0.25) & (absolute < 0.5), "target_magnitude/0.5<=abs<1": (absolute >= 0.5) & (absolute < 1.0), "target_magnitude/abs>=1": absolute >= 1.0, "cube_ownership/centered": context[:, 8] == 1, "cube_ownership/modeled_player": context[:, 9] == 1, "cube_ownership/opponent": context[:, 10] == 1, "crawford/false": context[:, 12] == 0, "crawford/true": context[:, 12] == 1, } cube = context[:, 6] for value in (1, 2, 4, 8): result[f"cube_value/{value}"] = cube == value result["cube_value/16+"] = cube >= 16 money, match_length = context[:, 0] == 1, context[:, 1] player_score, opponent_score = context[:, 2], context[:, 3] result.update({ "score/money": money, "score/0-0": (~money) & (player_score == 0) & (opponent_score == 0), "score/tied_nonzero": (~money) & (player_score == opponent_score) & (player_score > 0), "score/modeled_player_leads": (~money) & (player_score > opponent_score), "score/modeled_player_trails": (~money) & (player_score < opponent_score), "match_length/money": money, "match_length/1-3": (~money) & (match_length <= 3), "match_length/4-7": (match_length >= 4) & (match_length <= 7), "match_length/8-11": (match_length >= 8) & (match_length <= 11), "match_length/12+": match_length >= 12, }) for label in ("bar", "bearoff", "contact", "race"): result[f"position_class/{label}"] = classes == label return result class DetailedMetrics: def __init__(self) -> None: self.global_metric = RegressionMetrics() self.folds = {index: RegressionMetrics() for index in range(4)} self.segments: dict[str, RegressionMetrics] = {} def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray, masks: Mapping[str, np.ndarray]) -> None: self.global_metric.add(prediction, truth) for fold in range(4): selected = folds == fold if np.any(selected): self.folds[fold].add(prediction[selected], truth[selected]) for label, selected in masks.items(): if np.any(selected): self.segments.setdefault(label, RegressionMetrics()).add(prediction[selected], truth[selected]) def result(self) -> dict[str, Any]: return { "global": self.global_metric.result(), "development_complete_game_folds": { str(index): metric.result() for index, metric in self.folds.items() }, "segments": {label: metric.result() for label, metric in sorted(self.segments.items())}, } def _accepted_baseline(reference_models: Path): matches = [model for model in load_models(reference_models) if model.feature_set == "P3+CUBEFUL_CONTEXT" and model.checkpoint == "full"] if len(matches) != 1 or matches[0].alpha != 10.0: raise RuntimeError("unique accepted direct Cubeful baseline missing") return matches[0] def score_development( *, shallow_root: Path, split_manifest: Path, reference_models: Path, reference_development: Path, models_path: Path, output_path: Path, contributions_path: Path, ) -> dict[str, Any]: if not models_path.exists(): raise RuntimeError("models must be durable before DEVELOPMENT access") split = json.loads(split_manifest.read_text()) _, holdout, _ = _membership(split) campaign = str(split["source_authority"]["campaign"]) native_models = load_native_models(models_path) baseline = _accepted_baseline(reference_models) metrics = {"accepted_baseline": DetailedMetrics(), **{ model.model_id: DetailedMetrics() for model in native_models }} samples: tuple[np.ndarray, list[str], np.ndarray] | None = None rows = 0 decisions: set[str] = set() started = time.time() for path in _candidate_files(shallow_root): games = holdout.get(_partition_key(path), set()) if not games: exec /bin/bash -lc "jq '{package_identity_sha256,files}' artifacts/development/improve-additive-native-cubeful-modeling-v1/manifest.json jq '{identity_sha256, model_identities, partitions, target, protected_access_count}' results/improve-additive-native-cubeful-modeling-v1.json jq '{identity_sha256, package: .artifact_package_identity_sha256, terminal_classification, protected_access_count}' results/diagnose-native-cubeful-context-error-structure-v1.json jq '.models[] | select(.model_id==\"native-cubeful-p3-context-additive-ridge-v1\") | {model_id,model_identity_sha256,feature_ids:.transform.feature_ids, optimizer:{alpha:.optimizer.alpha,epochs:.optimizer.epochs,batch_size:.optimizer.batch_size,initial_learning_rate:.optimizer.initial_learning_rate,learning_rate_epoch_multiplier:.optimizer.learning_rate_epoch_multiplier,beta1:.optimizer.beta1,beta2:.optimizer.beta2,seed:.optimizer.seed,training_rows:.optimizer.training_rows,updates:.optimizer.updates}}' artifacts/development/improve-additive-native-cubeful-modeling-v1/models.json find ../runtime/native-cubeful-training-cache -maxdepth 2 -type f | wc -l du -sh ../runtime/native-cubeful-training-cache" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 206ms: { "package_identity_sha256": "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7", "files": [ { "path": "contribution-evidence.json", "sha256": "8bbf2ead035e69391c8c338993541a9ecd321b123c1a58aa575ff8677996dd02", "size_bytes": 130488 }, { "path": "development.json", "sha256": "906f25695ad3321b14b8b16da362dd23a7a5ec26c812022bdb2dfa8c1f81ab95", "size_bytes": 20130 }, { "path": "frozen-authorities.json", "sha256": "e2b36923b422505c7357eed2d346c5455a0566901587d73bd590efed3f123cbc", "size_bytes": 7616 }, { "path": "models.json", "sha256": "582e0eab87dc901c96f611287dec6eae9759995065ac4be633ff69f4ef6575cc", "size_bytes": 260053 }, { "path": "protected-access-log.json", "sha256": "8d18d2e10c46d4db96941f511b95717cf58d1186b547696a97c6f1281898538a", "size_bytes": 337 }, { "path": "protected-final.json", "sha256": "a9e429f40b93e9b7ba43359ccc53b959c6d9e14723e9ae687deff71975a52b75", "size_bytes": 335 }, { "path": "result-summary.json", "sha256": "3adbfd00470cfa2ad28c69ac24e9f980bb76fc1fb879639a3d325b41bc0b9468", "size_bytes": 8313 }, { "path": "test-evidence.json", "sha256": "dde02f2be47a681f7d8602457e38de43f5cea7e716d70ea79bcc7c7a1c5476ab", "size_bytes": 800 }, { "path": "training-cache.json", "sha256": "dd861f8c2a234947b477db28f429a3ea9bd8a42b379f26df65fc63fa3f714220", "size_bytes": 48695 }, { "path": "winner-freeze.json", "sha256": "ee6d5defca55012fd199d4c51c8af6bf88fbf17bb864591d34c790aa5265913c", "size_bytes": 1486 } ] } { "identity_sha256": "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be", "model_identities": { "native-cubeful-p3-additive-ridge-v1": "74a62b141204d65ecc70b50534cab260a7cca7ab08d63e4a927e8021e9b48044", "native-cubeful-p3-context-additive-ridge-v1": "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" }, "partitions": { "development": { "candidates": 2094039, "complete_games": 3178, "decisions": 100015, "membership_sha256": "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" }, "protected": { "access_before_frozen_winner": 0, "authority": "existing actual-4ply non-adaptive final evaluation", "candidates": 6963, "canonical_manifest_sha256": "effa2a8bc273be03222d8193c090f96ef3415224af228f0e45678b8e7ec498a7", "decisions": 2136, "split_assignment_sha256": "7dfb06ff31d6623bca3d3832a1e12b52e12bd1c991fa0db81c357d892b3302a8" }, "train": { "candidates": 20981224, "checkpoint": "1000000", "complete_games": 32228, "decisions": 1000002, "membership_sha256": "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" } }, "target": { "evaluation_mode": "Cubeful", "id": "native_cubeful_equity_static_next_player", "modeled_perspective": "normalized static post-move next player on roll", "source_field": "native_equity", "source_semantics": "GNU native checker-candidate Cubeful equity is maximized by the checker-move player", "transform": "-native_equity" }, "protected_access_count": 0 } { "identity_sha256": "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb", "package": null, "terminal_classification": "CONTEXT_FAILURE_LOCALIZED", "protected_access_count": 0 } { "model_id": "native-cubeful-p3-context-additive-ridge-v1", "model_identity_sha256": "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69", "feature_ids": [ "player_point_01_checkers", "player_point_02_checkers", "player_point_03_checkers", "player_point_04_checkers", "player_point_05_checkers", "player_point_06_checkers", "player_point_07_checkers", "player_point_08_checkers", "player_point_09_checkers", "player_point_10_checkers", "player_point_11_checkers", "player_point_12_checkers", "player_point_13_checkers", "player_point_14_checkers", "player_point_15_checkers", "player_point_16_checkers", "player_point_17_checkers", "player_point_18_checkers", "player_point_19_checkers", "player_point_20_checkers", "player_point_21_checkers", "player_point_22_checkers", "player_point_23_checkers", "player_point_24_checkers", "opponent_point_01_checkers", "opponent_point_02_checkers", "opponent_point_03_checkers", "opponent_point_04_checkers", "opponent_point_05_checkers", "opponent_point_06_checkers", "opponent_point_07_checkers", "opponent_point_08_checkers", "opponent_point_09_checkers", "opponent_point_10_checkers", "opponent_point_11_checkers", "opponent_point_12_checkers", "opponent_point_13_checkers", "opponent_point_14_checkers", "opponent_point_15_checkers", "opponent_point_16_checkers", "opponent_point_17_checkers", "opponent_point_18_checkers", "opponent_point_19_checkers", "opponent_point_20_checkers", "opponent_point_21_checkers", "opponent_point_22_checkers", "opponent_point_23_checkers", "opponent_point_24_checkers", "player_bar_checkers", "opponent_bar_checkers", "player_borne_off_checkers", "opponent_borne_off_checkers", "player_point_01_blot", "player_point_01_made", "player_point_01_spares", "player_point_01_stack_over_4", "player_point_02_blot", "player_point_02_made", "player_point_02_spares", "player_point_02_stack_over_4", "player_point_03_blot", "player_point_03_made", "player_point_03_spares", "player_point_03_stack_over_4", "player_point_04_blot", "player_point_04_made", "player_point_04_spares", "player_point_04_stack_over_4", "player_point_05_blot", "player_point_05_made", "player_point_05_spares", "player_point_05_stack_over_4", "player_point_06_blot", "player_point_06_made", "player_point_06_spares", "player_point_06_stack_over_4", "player_point_07_blot", "player_point_07_made", "player_point_07_spares", "player_point_07_stack_over_4", "player_point_08_blot", "player_point_08_made", "player_point_08_spares", "player_point_08_stack_over_4", "player_point_09_blot", "player_point_09_made", "player_point_09_spares", "player_point_09_stack_over_4", "player_point_10_blot", "player_point_10_made", "player_point_10_spares", "player_point_10_stack_over_4", "player_point_11_blot", "player_point_11_made", "player_point_11_spares", "player_point_11_stack_over_4", "player_point_12_blot", "player_point_12_made", "player_point_12_spares", "player_point_12_stack_over_4", "player_point_13_blot", "player_point_13_made", "player_point_13_spares", "player_point_13_stack_over_4", "player_point_14_blot", "player_point_14_made", "player_point_14_spares", "player_point_14_stack_over_4", "player_point_15_blot", "player_point_15_made", "player_point_15_spares", "player_point_15_stack_over_4", "player_point_16_blot", "player_point_16_made", "player_point_16_spares", "player_point_16_stack_over_4", "player_point_17_blot", "player_point_17_made", "player_point_17_spares", "player_point_17_stack_over_4", "player_point_18_blot", "player_point_18_made", "player_point_18_spares", "player_point_18_stack_over_4", "player_point_19_blot", "player_point_19_made", "player_point_19_spares", "player_point_19_stack_over_4", "player_point_20_blot", "player_point_20_made", "player_point_20_spares", "player_point_20_stack_over_4", "player_point_21_blot", "player_point_21_made", "player_point_21_spares", "player_point_21_stack_over_4", "player_point_22_blot", "player_point_22_made", "player_point_22_spares", "player_point_22_stack_over_4", "player_point_23_blot", "player_point_23_made", "player_point_23_spares", "player_point_23_stack_over_4", "player_point_24_blot", "player_point_24_made", "player_point_24_spares", "player_point_24_stack_over_4", "opponent_point_01_blot", "opponent_point_01_made", "opponent_point_01_spares", "opponent_point_01_stack_over_4", "opponent_point_02_blot", "opponent_point_02_made", "opponent_point_02_spares", "opponent_point_02_stack_over_4", "opponent_point_03_blot", "opponent_point_03_made", "opponent_point_03_spares", "opponent_point_03_stack_over_4", "opponent_point_04_blot", "opponent_point_04_made", "opponent_point_04_spares", "opponent_point_04_stack_over_4", "opponent_point_05_blot", "opponent_point_05_made", "opponent_point_05_spares", "opponent_point_05_stack_over_4", "opponent_point_06_blot", "opponent_point_06_made", "opponent_point_06_spares", "opponent_point_06_stack_over_4", "opponent_point_07_blot", "opponent_point_07_made", "opponent_point_07_spares", "opponent_point_07_stack_over_4", "opponent_point_08_blot", "opponent_point_08_made", "opponent_point_08_spares", "opponent_point_08_stack_over_4", "opponent_point_09_blot", "opponent_point_09_made", "opponent_point_09_spares", "opponent_point_09_stack_over_4", "opponent_point_10_blot", "opponent_point_10_made", "opponent_point_10_spares", "opponent_point_10_stack_over_4", "opponent_point_11_blot", "opponent_point_11_made", "opponent_point_11_spares", "opponent_point_11_stack_over_4", "opponent_point_12_blot", "opponent_point_12_made", "opponent_point_12_spares", "opponent_point_12_stack_over_4", "opponent_point_13_blot", "opponent_point_13_made", "opponent_point_13_spares", "opponent_point_13_stack_over_4", "opponent_point_14_blot", "opponent_point_14_made", "opponent_point_14_spares", "opponent_point_14_stack_over_4", "opponent_point_15_blot", "opponent_point_15_made", "opponent_point_15_spares", "opponent_point_15_stack_over_4", "opponent_point_16_blot", "opponent_point_16_made", "opponent_point_16_spares", "opponent_point_16_stack_over_4", "opponent_point_17_blot", "opponent_point_17_made", "opponent_point_17_spares", "opponent_point_17_stack_over_4", "opponent_point_18_blot", "opponent_point_18_made", "opponent_point_18_spares", "opponent_point_18_stack_over_4", "opponent_point_19_blot", "opponent_point_19_made", "opponent_point_19_spares", "opponent_point_19_stack_over_4", "opponent_point_20_blot", "opponent_point_20_made", "opponent_point_20_spares", "opponent_point_20_stack_over_4", "opponent_point_21_blot", "opponent_point_21_made", "opponent_point_21_spares", "opponent_point_21_stack_over_4", "opponent_point_22_blot", "opponent_point_22_made", "opponent_point_22_spares", "opponent_point_22_stack_over_4", "opponent_point_23_blot", "opponent_point_23_made", "opponent_point_23_spares", "opponent_point_23_stack_over_4", "opponent_point_24_blot", "opponent_point_24_made", "opponent_point_24_spares", "opponent_point_24_stack_over_4", "player_pip_count", "opponent_pip_count", "relative_pip_difference", "player_rearmost_point", "player_made_home_points", "opponent_made_home_points", "player_blot_count", "opponent_blot_count", "player_direct_hit_die_count", "opponent_entry_failure_probability", "player_longest_prime", "opponent_longest_prime", "player_anchor_count", "opponent_anchor_count", "player_occupied_points", "player_spare_checkers", "player_max_stack", "player_home_board_checkers", "opponent_home_board_checkers", "player_outer_board_checkers", "opponent_outer_board_checkers", "player_mid_board_checkers", "opponent_mid_board_checkers", "player_far_board_checkers", "opponent_far_board_checkers", "opponent_made_outer_points", "player_made_mid_points", "opponent_made_mid_points", "player_made_far_points", "opponent_made_far_points", "player_made_points", "opponent_made_points", "player_home_blots", "opponent_home_blots", "player_outer_blots", "opponent_outer_blots", "player_mid_blots", "opponent_mid_blots", "player_far_blots", "opponent_far_blots", "player_home_occupied_points", "opponent_home_occupied_points", "player_outer_occupied_points", "opponent_outer_occupied_points", "player_mid_occupied_points", "opponent_mid_occupied_points", "player_far_occupied_points", "opponent_far_occupied_points", "opponent_occupied_points", "opponent_spare_checkers", "opponent_max_stack", "player_stack_excess_square", "opponent_stack_excess_square", "player_stack_square_sum", "opponent_stack_square_sum", "player_checker_point_mean", "opponent_checker_point_mean", "player_checker_point_variance", "opponent_checker_point_variance", "player_occupied_point_span", "opponent_occupied_point_span", "player_made_point_span", "opponent_made_point_span", "opponent_rearmost_point", "player_frontmost_point", "opponent_frontmost_point", "player_home_longest_prime", "opponent_home_longest_prime", "player_outer_longest_prime", "opponent_outer_longest_prime", "contact_overlap_distance", "player_checker_point_mean_absolute_deviation", "player_checker_point_standard_deviation", "player_checker_point_skewness", "player_checker_point_q25", "player_checker_point_q50", "player_checker_point_q75", "opponent_checker_point_mean_absolute_deviation", "opponent_checker_point_standard_deviation", "opponent_checker_point_skewness", "opponent_checker_point_q25", "opponent_checker_point_q50", "opponent_checker_point_q75", "player_made_window_count_length_2", "player_made_window_count_length_3", "player_made_window_count_length_4", "player_made_window_count_length_5", "player_made_window_count_length_6", "opponent_made_window_count_length_2", "opponent_made_window_count_length_3", "opponent_made_window_count_length_4", "opponent_made_window_count_length_5", "opponent_made_window_count_length_6", "player_point_count_at_least_3_checkers", "player_point_count_at_least_4_checkers", "player_point_count_at_least_5_checkers", "player_point_count_at_least_6_checkers", "opponent_point_count_at_least_3_checkers", "opponent_point_count_at_least_4_checkers", "opponent_point_count_at_least_5_checkers", "opponent_point_count_at_least_6_checkers", "player_made_point_center", "player_made_point_mean_absolute_deviation", "player_made_point_longest_gap", "opponent_made_point_center", "opponent_made_point_mean_absolute_deviation", "opponent_made_point_longest_gap", "cubeful_is_money", "cubeful_match_length", "cubeful_player_score", "cubeful_opponent_score", "cubeful_player_away", "cubeful_opponent_away", "cubeful_cube_value", "cubeful_cube_log2", "cubeful_cube_centered", "cubeful_cube_owned_by_player", "cubeful_cube_owned_by_opponent", "cubeful_cube_owner_relative_code", "cubeful_crawford", "cubeful_jacoby", "cubeful_cube_offer_pending" ], "optimizer": { "alpha": 100.0, "epochs": 12, "batch_size": 32768, "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, "beta1": 0.9, "beta2": 0.999, "seed": 20260823, "training_rows": 20981224, "updates": 8220 } } 246 5.9G ../runtime/native-cubeful-training-cache exec /bin/bash -lc "jq 'keys' results/diagnose-native-cubeful-context-error-structure-v1.json jq '{identity_sha256,artifact_package_identity_sha256,package_identity_sha256,terminal_classification,protected_access_count}' results/diagnose-native-cubeful-context-error-structure-v1.json sha256sum artifacts/development/improve-additive-native-cubeful-modeling-v1/* | sort sha256sum artifacts/development/diagnose-native-cubeful-context-error-structure-v1/* | sort | sed -n '1,40p' sha256sum ../tm/milestones/explainer-native-cubeful-k001/prompts/001-improve-additive-native-cubeful-modeling-v1.md ../tm/milestones/explainer-native-cubeful-k001/prompts/002-diagnose-native-cubeful-context-error-structure-v1.md ../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 100ms: [ "accepted_product_architecture", "analyzer", "calculated_cubeful", "candidate_disposition", "canonical", "corpus", "development_membership_identity_sha256", "generation_1_package_identity_sha256", "generation_1_result_identity_sha256", "identity_sha256", "model_fits", "modeling_candidates", "next_task_status", "prediction_stream_logical_identities", "production", "protected_access_count", "routing_cell", "stable_context_regression_cells", "status", "terminal_classification", "version" ] { "identity_sha256": "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb", "artifact_package_identity_sha256": null, "package_identity_sha256": null, "terminal_classification": "CONTEXT_FAILURE_LOCALIZED", "protected_access_count": 0 } 3adbfd00470cfa2ad28c69ac24e9f980bb76fc1fb879639a3d325b41bc0b9468 artifacts/development/improve-additive-native-cubeful-modeling-v1/result-summary.json 41872900238dcb836cfde503e127f9d7da4f13ba211f5411dfa92b52a5350bff artifacts/development/improve-additive-native-cubeful-modeling-v1/self-verification.json 582e0eab87dc901c96f611287dec6eae9759995065ac4be633ff69f4ef6575cc artifacts/development/improve-additive-native-cubeful-modeling-v1/models.json 86728c4ee77a3f21afaa761709b06eaed973d063b3989f034b6ea67c5bc279e7 artifacts/development/improve-additive-native-cubeful-modeling-v1/manifest.json 8bbf2ead035e69391c8c338993541a9ecd321b123c1a58aa575ff8677996dd02 artifacts/development/improve-additive-native-cubeful-modeling-v1/contribution-evidence.json 8d18d2e10c46d4db96941f511b95717cf58d1186b547696a97c6f1281898538a artifacts/development/improve-additive-native-cubeful-modeling-v1/protected-access-log.json 906f25695ad3321b14b8b16da362dd23a7a5ec26c812022bdb2dfa8c1f81ab95 artifacts/development/improve-additive-native-cubeful-modeling-v1/development.json a9e429f40b93e9b7ba43359ccc53b959c6d9e14723e9ae687deff71975a52b75 artifacts/development/improve-additive-native-cubeful-modeling-v1/protected-final.json b96329eb5764bee4d2552ac6a15b86810fef7ea979a08c07fb2b695c0b24b328 artifacts/development/improve-additive-native-cubeful-modeling-v1/SHA256SUMS dd861f8c2a234947b477db28f429a3ea9bd8a42b379f26df65fc63fa3f714220 artifacts/development/improve-additive-native-cubeful-modeling-v1/training-cache.json dde02f2be47a681f7d8602457e38de43f5cea7e716d70ea79bcc7c7a1c5476ab artifacts/development/improve-additive-native-cubeful-modeling-v1/test-evidence.json e077e5c24d523a98c1faedbc5873ceee9286720a8f557a4087e655cff21c3378 artifacts/development/improve-additive-native-cubeful-modeling-v1/preflight-log.jsonl e2b36923b422505c7357eed2d346c5455a0566901587d73bd590efed3f123cbc artifacts/development/improve-additive-native-cubeful-modeling-v1/frozen-authorities.json ee6d5defca55012fd199d4c51c8af6bf88fbf17bb864591d34c790aa5265913c artifacts/development/improve-additive-native-cubeful-modeling-v1/winner-freeze.json 069da60d9939bbf163d50139093b835fc1b237c0df4d354346a847e3e88df881 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/test-evidence.json 1ac645a1755cc419d3ed29e480d3210084c928d6935ace52329d1e6b97b48f39 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/manifest.json 1e5c20f940765f2801f42ad477bcddcf68130d6a7238c68fc0bce6a7ef80eb99 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/SHA256SUMS 2aca347084fa2f15105629d2c60d16f899c18552fe0bf44b5918949e8a099837 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/preflight-log.jsonl 3e5279842536728291eade3c8c5fcab40f03fee9454f510c38fbce1690ca00c8 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/run-record.json 443976e5b9d53ce70d28ff19b894cd7850a72d5565a4ccc486faadc14270cbb8 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/cell-metrics.json 4d1869c777a421fea6a1aee9ac5673470b80be513a8abd3e2ab86afabd0f3b70 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/generation-1-reproduction.json 54d9bdf9b233eee82e079e9e2f028a840aae06df6206489065df68d32e5e530d artifacts/development/diagnose-native-cubeful-context-error-structure-v1/input-receipt.json 5d385dbe1461f41a28bc9652f894d1dc2a94384b785f1245a3519fcb3575837d artifacts/development/diagnose-native-cubeful-context-error-structure-v1/prediction-stream-identities.json 62e95b3ba50455f00990a6e9dcfa90fca718cd6bf0d5ef6739e4484c58508de5 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/segment-membership.json 6422dcaa74b38037d0938c8cbe9b73675125859d6029420725b64e46622496c6 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/deterministic-verification.json 683b75e325a10724fa51b98733d34901d0f946bef25d31ab7ab855c643f0dfa0 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/self-verification.json 6c77b74d764cfb5231eb0e2e3678911568d5a6b205c833a6fe6ab1ea68dfed6c artifacts/development/diagnose-native-cubeful-context-error-structure-v1/result-summary.json 8fc1a51e8d122079c90a0fdb00c4fb2c0355e47a71a0809ca70580641267e1d3 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/fold-evidence.json 98b79e5c39d569d66b92e714035c8627f0b6cbbca76d689d71de2809b9b8a90c artifacts/development/diagnose-native-cubeful-context-error-structure-v1/protected-access-log.json b85ffbf1ca86c2c23c02c8199ffd6e5b84287499cd6e68db30f5ca101ecdb166 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/field-mapping-receipt.json c4d9f79180d7f0b085b208ab40e2d144710a7e4ad238ed03138a46eda8e41b87 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/regression-mass.json e290d584852e3a9d16040379620f69a803d838d6927fe576d1fa194a2be3d603 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/support-stability.json f5f818c25036865e626b858e453b1943d86e643ffbf24f809350d369703f3206 artifacts/development/diagnose-native-cubeful-context-error-structure-v1/routing-record.json debfc2521be4c58126308303e2fe4a405d8885ab5db0f1258e1e83e4abd6662a ../tm/milestones/explainer-native-cubeful-k001/prompts/001-improve-additive-native-cubeful-modeling-v1.md 36ff14d5004357deb7d1e57f0b416e408ac93cae2297a600114562fc82a95a69 ../tm/milestones/explainer-native-cubeful-k001/prompts/002-diagnose-native-cubeful-context-error-structure-v1.md 740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4 ../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md exec /bin/bash -lc "jq '{package_identity_sha256,files: [.files[]|.path]}' artifacts/development/diagnose-native-cubeful-context-error-structure-v1/manifest.json jq '{version,status,identity_sha256,terminal_classification,routing_cell,generation_1_result_identity_sha256,generation_1_package_identity_sha256,protected_access_count}' results/diagnose-native-cubeful-context-error-structure-v1.json jq '{identity_sha256,source_and_partition_authority,target_authority,feature_authority,model_comparison,development_grouped_folds,host_policy,activity_boundary}' artifacts/development/improve-additive-native-cubeful-modeling-v1/frozen-authorities.json > /tmp/native-authorities-slice.json sed -n '1,320p' /tmp/native-authorities-slice.json" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 65ms: { "package_identity_sha256": "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7", "files": [ "cell-metrics.json", "deterministic-verification.json", "field-mapping-receipt.json", "fold-evidence.json", "generation-1-reproduction.json", "input-receipt.json", "prediction-stream-identities.json", "preflight-log.jsonl", "protected-access-log.json", "regression-mass.json", "result-summary.json", "routing-record.json", "run-record.json", "segment-membership.json", "support-stability.json", "test-evidence.json" ] } { "version": "diagnose-native-cubeful-context-error-structure-v1-result-v1", "status": "PASS_COMPLETE", "identity_sha256": "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb", "terminal_classification": "CONTEXT_FAILURE_LOCALIZED", "routing_cell": { "cell": "money_or_zero", "family": "crawford_state" }, "generation_1_result_identity_sha256": "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be", "generation_1_package_identity_sha256": "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7", "protected_access_count": 0 } { "identity_sha256": "395b80cf6a1ec85f3e76745dd1430fec4a82520ecbb0224554b4bbcc9b153c9a", "source_and_partition_authority": { "development": { "candidates": 2094039, "complete_games": 3178, "decisions": 100015, "membership_sha256": "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" }, "excluded_decision_membership_sha256": "c114765ae3e48330d928d665dbfdff4e902e90f2b44e0e1577006b35a161b6b9", "protected": { "access_before_frozen_winner": 0, "authority": "existing actual-4ply non-adaptive final evaluation", "candidates": 6963, "canonical_manifest_sha256": "effa2a8bc273be03222d8193c090f96ef3415224af228f0e45678b8e7ec498a7", "decisions": 2136, "split_assignment_sha256": "7dfb06ff31d6623bca3d3832a1e12b52e12bd1c991fa0db81c357d892b3302a8" }, "source": "accepted GNU 0-ply modeling rows; no new engine work", "source_manifest_sha256": "a756567c6c6e316f0bf45527e516127af4317ab8e3b9ea5aa0156ca2dae1c14d", "split_manifest_identity_sha256": "eb0d571182b529588861731d29a49c26880d93f0170646b16e8974d81576ed0f", "split_manifest_sha256": "1d125f02d9c5c7340528e134ee6e2815d3ebe612e9df009ae236b726a92019d6", "train": { "candidates": 20981224, "checkpoint": "1000000", "complete_games": 32228, "decisions": 1000002, "membership_sha256": "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" } }, "target_authority": { "evaluation_mode": "Cubeful", "id": "native_cubeful_equity_static_next_player", "modeled_perspective": "normalized static post-move next player on roll", "source_field": "native_equity", "source_semantics": "GNU native checker-candidate Cubeful equity is maximized by the checker-move player", "transform": "-native_equity" }, "feature_authority": { "context_feature_count": 15, "context_fields": [ "cubeful_is_money", "cubeful_match_length", "cubeful_player_score", "cubeful_opponent_score", "cubeful_player_away", "cubeful_opponent_away", "cubeful_cube_value", "cubeful_cube_log2", "cubeful_cube_centered", "cubeful_cube_owned_by_player", "cubeful_cube_owned_by_opponent", "cubeful_cube_owner_relative_code", "cubeful_crawford", "cubeful_jacoby", "cubeful_cube_offer_pending" ], "context_order_sha256": "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914", "direct_registry_sha256": "b664347c564c0fecd53448554119941f94bb161075a61bf6cb957145949e604e", "position_feature_count": 351, "position_registry": "explainer-position-value-p3-v1-30ede35745bbbc64", "position_registry_sha256": "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" }, "model_comparison": { "additive_basis": { "hinge_quantiles": [ 0.25, 0.5, 0.75 ], "interactions": 0, "intercept": true, "knots": "TRAIN-only", "per_feature_terms": [ "standardized_linear", "hinge_q25", "hinge_q50", "hinge_q75" ] }, "baseline": { "accepted_development_evidence_sha256": "29d09310c078788685d653fb4b16e8fef84959d5ca1f6c002baad6f4eb532091", "alpha": 10.0, "feature_count": 366, "id": "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1", "model_identity_sha256": "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5", "models_artifact_sha256": "b27779b1947bcae91bddb067bd03d2267016140839f5b2bf9ce793ce6797c493", "training_checkpoint": "full" }, "candidates": [ { "feature_count": 351, "features": "P3 position only", "id": "native-cubeful-p3-additive-ridge-v1" }, { "feature_count": 366, "features": "P3 plus the complete accepted 15-field factual context block", "id": "native-cubeful-p3-context-additive-ridge-v1" } ], "optimizer": { "algorithm": "deterministic streaming Adam on the Ridge objective", "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, "epochs": 12, "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, "training_arithmetic": "float32 basis/optimizer; float64 retained inference and reconstruction" }, "regularization_disposition": "singleton reuse of the accepted successful ADDEQ alpha; no outcome search", "regularization_grid": [ 100.0 ] }, "development_grouped_folds": { "assignment": "SHA256([version,seed,complete_game_id]) modulo four", "folds": 4, "seed": 20260823, "version": "native-cubeful-development-complete-game-fold-4-v1" }, "host_policy": { "blas_threads": 1, "fresh_preflight_before_each_substantial_phase": true, "host": "high-fructose-corn-syrup", "nice": 19, "protected_process_intervention": "FORBIDDEN", "workers": 1 }, "activity_boundary": { "analyzer_mutation": false, "canonical_mutation": false, "corpus_mutation": false, "new_generic_0ply": 0, "new_gnu": 0, "new_labels": 0, "new_matches": 0, "new_sage": 0, "production_promotion": false, "sage_gnu_campaign_training_rows": 0 } } exec /bin/bash -lc "PYTHONPATH=src ../runtime/python311/usr/bin/python3 -c 'from backgammon_explainer.native_cubeful_experiment import _sha; fields=[\"cubeful_is_money\",\"cubeful_match_length\",\"cubeful_player_score\",\"cubeful_opponent_score\",\"cubeful_player_away\",\"cubeful_opponent_away\",\"cubeful_cube_value\",\"cubeful_cube_log2\",\"cubeful_cube_centered\",\"cubeful_cube_owned_by_player\",\"cubeful_cube_owned_by_opponent\",\"cubeful_cube_owner_relative_code\",\"cubeful_jacoby\",\"cubeful_cube_offer_pending\"]; print(_sha(fields))'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl exited 127 in 31ms: /bin/bash: line 1: ../runtime/python311/usr/bin/python3: No such file or directory exec /bin/bash -lc "rg --files ../runtime/python311 | rg '/(python3|python)"'$|pyvenv.cfg'"' | sed -n '1,40p' command -v python3 python3 --version" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 74ms: /usr/bin/python3 Python 3.13.5 exec /bin/bash -lc "rg --files -uu . | rg '("'^|/)(python|python3|pytest)$|pyvenv.cfg|uv.lock|pyproject.toml'"' | sed -n '1,100p' ls -la ../runtime/python311 find ../runtime/python311 -maxdepth 3 -type f -name 'python*' -o -type l -name 'python*' | sed -n '1,50p'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 88ms: ./pyproject.toml total 17 drwxr-xr-x 4 22058 22058 4 Jan 31 2023 . drwxr-xr-x 8 22058 22058 16 Aug 24 04:09 .. drwxr-xr-x 3 22058 22058 3 Mar 13 2023 etc drwxr-xr-x 6 22058 22058 6 Jan 31 2023 usr ../runtime/python311/usr/bin/python3.11 exec /bin/bash -lc 'ls -l ../runtime/python311/usr/bin/python3.11 file ../runtime/python311/usr/bin/python3.11 ldd ../runtime/python311/usr/bin/python3.11 || true ls -l /lib64/ld-linux-x86-64.so.2 /lib/x86_64-linux-gnu/ld-linux-x86-64.so.2 2>/dev/null || true' in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 150ms: -rwxr-xr-x 1 22058 22058 6831704 Mar 13 2023 ../runtime/python311/usr/bin/python3.11 ../runtime/python311/usr/bin/python3.11: ELF 64-bit LSB executable, x86-64, version 1 (SYSV), dynamically linked, interpreter /lib64/ld-linux-x86-64.so.2, BuildID[sha1]=ac175ec7666754cf818b271b4fdc2761ac6865f2, for GNU/Linux 3.2.0, stripped linux-vdso.so.1 (0x00007f03c2c8b000) libm.so.6 => /lib/x86_64-linux-gnu/libm.so.6 (0x00007f03c2b75000) libz.so.1 => /lib/x86_64-linux-gnu/libz.so.1 (0x00007f03c2b55000) libexpat.so.1 => /lib/x86_64-linux-gnu/libexpat.so.1 (0x00007f03c2b25000) libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6 (0x00007f03c292d000) /lib64/ld-linux-x86-64.so.2 (0x00007f03c2c8d000) -rwxr-xr-x 1 nobody nogroup 225672 Apr 27 16:09 /lib/x86_64-linux-gnu/ld-linux-x86-64.so.2 lrwxrwxrwx 1 nobody nogroup 44 Apr 27 16:09 /lib64/ld-linux-x86-64.so.2 -> ../lib/x86_64-linux-gnu/ld-linux-x86-64.so.2 exec /bin/bash -lc "PYTHONHOME=../runtime/python311 ../runtime/python311/usr/bin/python3.11 --version PYTHONHOME=../runtime/python311 PYTHONPATH=src ../runtime/python311/usr/bin/python3.11 -c 'import numpy,pyarrow,sklearn; print(numpy.__version__, pyarrow.__version__, sklearn.__version__)'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl exited 1 in 67ms: Python 3.11.2 Python path configuration: PYTHONHOME = '../runtime/python311' PYTHONPATH = 'src' program name = '../runtime/python311/usr/bin/python3.11' isolated = 0 environment = 1 user site = 1 safe_path = 0 import site = 1 is in build tree = 0 stdlib dir = '../runtime/python311/lib/python3.11' sys._base_executable = '/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/../runtime/python311/usr/bin/python3.11' sys.base_prefix = '../runtime/python311' sys.base_exec_prefix = '../runtime/python311' sys.platlibdir = 'lib' sys.executable = '/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/../runtime/python311/usr/bin/python3.11' sys.prefix = '../runtime/python311' sys.exec_prefix = '../runtime/python311' sys.path = [ '/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src', '../runtime/python311/lib/python311.zip', '../runtime/python311/lib/python3.11', '../runtime/python311/lib/python3.11/lib-dynload', ] Fatal Python error: init_fs_encoding: failed to get the Python codec of the filesystem encoding Python runtime state: core initialized ModuleNotFoundError: No module named 'encodings' Current thread 0x00007f0c448c6740 (most recent call first): exec /bin/bash -lc "PYTHONHOME=../runtime/python311/usr PYTHONPATH=src ../runtime/python311/usr/bin/python3.11 -c 'import sys,numpy,pyarrow,sklearn; print(sys.version); print(numpy.__version__, pyarrow.__version__, sklearn.__version__)'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl exited 1 in 82ms: Traceback (most recent call last): File "", line 1, in ModuleNotFoundError: No module named 'numpy' exec /bin/bash -lc "find ../runtime -maxdepth 5 -type d \\( -name site-packages -o -name dist-packages \\) -print python3 -c 'import numpy,pyarrow,sklearn; print(numpy.__version__,pyarrow.__version__,sklearn.__version__)'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl exited 1 in 158ms: ../runtime/uv-failed-venv-20260824T0312/lib/python3.13/site-packages ../runtime/python311/usr/lib/python3/dist-packages Traceback (most recent call last): File "", line 1, in import numpy,pyarrow,sklearn; print(numpy.__version__,pyarrow.__version__,sklearn.__version__) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ModuleNotFoundError: No module named 'numpy' exec /bin/bash -lc "find /users/a2andrad -path '*/site-packages/numpy' -type d -print 2>/dev/null | sed -n '1,30p' find /tmp -path '*/site-packages/numpy' -type d -print 2>/dev/null | sed -n '1,30p' command -v uv || true uv --version 2>/dev/null || true" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl exec /bin/bash -lc "find ../runtime -maxdepth 3 -mindepth 1 -type d -printf '%p\\n' | sed -n '1,200p' find ../runtime -type d -name 'numpy*' -o -type d -name 'pyarrow*' | sed -n '1,100p' find /tmp -maxdepth 4 -type d -name 'numpy*' -o -type d -name 'pyarrow*' | sed -n '1,100p'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 191ms: ../runtime/native-cubeful-training-cache ../runtime/native-cubeful-training-cache/partition-026 ../runtime/native-cubeful-training-cache/partition-051 ../runtime/native-cubeful-training-cache/partition-063 ../runtime/native-cubeful-training-cache/partition-014 ../runtime/native-cubeful-training-cache/partition-069 ../runtime/native-cubeful-training-cache/partition-048 ../runtime/native-cubeful-training-cache/partition-035 ../runtime/native-cubeful-training-cache/partition-042 ../runtime/native-cubeful-training-cache/partition-070 ../runtime/native-cubeful-training-cache/partition-007 ../runtime/native-cubeful-training-cache/partition-000 ../runtime/native-cubeful-training-cache/partition-077 ../runtime/native-cubeful-training-cache/partition-038 ../runtime/native-cubeful-training-cache/partition-045 ../runtime/native-cubeful-training-cache/partition-032 ../runtime/native-cubeful-training-cache/partition-013 ../runtime/native-cubeful-training-cache/partition-064 ../runtime/native-cubeful-training-cache/partition-019 ../runtime/native-cubeful-training-cache/partition-080 ../runtime/native-cubeful-training-cache/partition-056 ../runtime/native-cubeful-training-cache/partition-021 ../runtime/native-cubeful-training-cache/partition-058 ../runtime/native-cubeful-training-cache/partition-052 ../runtime/native-cubeful-training-cache/partition-025 ../runtime/native-cubeful-training-cache/partition-017 ../runtime/native-cubeful-training-cache/partition-060 ../runtime/native-cubeful-training-cache/partition-041 ../runtime/native-cubeful-training-cache/partition-036 ../runtime/native-cubeful-training-cache/partition-004 ../runtime/native-cubeful-training-cache/partition-073 ../runtime/native-cubeful-training-cache/partition-079 ../runtime/native-cubeful-training-cache/partition-074 ../runtime/native-cubeful-training-cache/partition-003 ../runtime/native-cubeful-training-cache/partition-009 ../runtime/native-cubeful-training-cache/partition-031 ../runtime/native-cubeful-training-cache/partition-046 ../runtime/native-cubeful-training-cache/partition-067 ../runtime/native-cubeful-training-cache/partition-010 ../runtime/native-cubeful-training-cache/partition-028 ../runtime/native-cubeful-training-cache/partition-022 ../runtime/native-cubeful-training-cache/partition-055 ../runtime/native-cubeful-training-cache/partition-018 ../runtime/native-cubeful-training-cache/partition-081 ../runtime/native-cubeful-training-cache/partition-065 ../runtime/native-cubeful-training-cache/partition-012 ../runtime/native-cubeful-training-cache/partition-020 ../runtime/native-cubeful-training-cache/partition-057 ../runtime/native-cubeful-training-cache/partition-076 ../runtime/native-cubeful-training-cache/partition-001 ../runtime/native-cubeful-training-cache/partition-033 ../runtime/native-cubeful-training-cache/partition-044 ../runtime/native-cubeful-training-cache/partition-039 ../runtime/native-cubeful-training-cache/partition-043 ../runtime/native-cubeful-training-cache/partition-034 ../runtime/native-cubeful-training-cache/partition-049 ../runtime/native-cubeful-training-cache/partition-006 ../runtime/native-cubeful-training-cache/partition-071 ../runtime/native-cubeful-training-cache/partition-050 ../runtime/native-cubeful-training-cache/partition-027 ../runtime/native-cubeful-training-cache/partition-068 ../runtime/native-cubeful-training-cache/partition-015 ../runtime/native-cubeful-training-cache/partition-062 ../runtime/native-cubeful-training-cache/partition-011 ../runtime/native-cubeful-training-cache/partition-066 ../runtime/native-cubeful-training-cache/partition-054 ../runtime/native-cubeful-training-cache/partition-023 ../runtime/native-cubeful-training-cache/partition-029 ../runtime/native-cubeful-training-cache/partition-008 ../runtime/native-cubeful-training-cache/partition-002 ../runtime/native-cubeful-training-cache/partition-075 ../runtime/native-cubeful-training-cache/partition-047 ../runtime/native-cubeful-training-cache/partition-030 ../runtime/native-cubeful-training-cache/partition-037 ../runtime/native-cubeful-training-cache/partition-040 ../runtime/native-cubeful-training-cache/partition-078 ../runtime/native-cubeful-training-cache/partition-072 ../runtime/native-cubeful-training-cache/partition-005 ../runtime/native-cubeful-training-cache/partition-024 ../runtime/native-cubeful-training-cache/partition-053 ../runtime/native-cubeful-training-cache/partition-059 ../runtime/native-cubeful-training-cache/partition-061 ../runtime/native-cubeful-training-cache/partition-016 ../runtime/uv-failed-venv-20260824T0312 ../runtime/uv-failed-venv-20260824T0312/bin ../runtime/uv-failed-venv-20260824T0312/lib ../runtime/uv-failed-venv-20260824T0312/lib/python3.13 ../runtime/uv-cache ../runtime/uv-cache/interpreter-v4 ../runtime/uv-cache/interpreter-v4/501209b0da452948 ../runtime/uv-cache/sdists-v9 ../runtime/uv-cache/sdists-v9/editable ../runtime/uv-cache/builds-v0 ../runtime/python311 ../runtime/python311/usr ../runtime/python311/usr/bin ../runtime/python311/usr/share ../runtime/python311/usr/lib ../runtime/python311/usr/include ../runtime/python311/etc ../runtime/python311/etc/python3.11 ../runtime/uv-python ../runtime/uv-python/.temp ../runtime/venv find: ‘/tmp/systemd-private-4bf5c6bf2f174d558e9778cbfe5fd4f7-munin-node.service-SIOpRd’: Permission denied find: ‘/tmp/tmux-25924’: Permission denied find: ‘/tmp/nvim.j353song’: Permission denied find: ‘/tmp/tmux-26109’: Permission denied find: ‘/tmp/forge_installer2447184155977020961’: Permission denied find: ‘/tmp/systemd-private-4bf5c6bf2f174d558e9778cbfe5fd4f7-systemd-logind.service-fDgKbb’: Permission denied find: ‘/tmp/nix-shell.81nYRT’: Permission denied find: ‘/tmp/nix-3736550-1042777744/scratch-reserved’: Permission denied find: ‘/tmp/nix-3736550-1042777744/u’: Permission denied find: ‘/tmp/nix-3736550-1042777744/tmp’: Permission denied find: ‘/tmp/nix-3736550-1042777744/music’: Permission denied find: ‘/tmp/nix-3736550-1042777744/scratch’: Permission denied find: ‘/tmp/nix-3736550-1042777744/home’: Permission denied find: ‘/tmp/nix-3736550-1042777744/dev’: Permission denied find: ‘/tmp/nix-3736550-1042777744/scratch-hfcs’: Permission denied find: ‘/tmp/nix-3736550-1042777744/srv’: Permission denied find: ‘/tmp/nix-3736550-1042777744/usr’: Permission denied find: ‘/tmp/nix-3736550-1042777744/mnt’: Permission denied find: ‘/tmp/nix-3736550-1042777744/run’: Permission denied find: ‘/tmp/nix-3736550-1042777744/users’: Permission denied find: ‘/tmp/nix-3736550-1042777744/media’: Permission denied find: ‘/tmp/nix-3736550-1042777744/opt’: Permission denied find: ‘/tmp/nix-3736550-1042777744/root’: Permission denied find: ‘/tmp/nix-3736550-1042777744/sys’: Permission denied find: ‘/tmp/nix-3736550-1042777744/var’: Permission denied find: ‘/tmp/nix-3736550-1042777744/lost+found’: Permission denied find: ‘/tmp/nix-3736550-1042777744/proc’: Permission denied find: ‘/tmp/nix-3736550-1042777744/boot’: Permission denied find: ‘/tmp/nix-3736550-1042777744/etc’: Permission denied find: ‘/tmp/nix-shell.563izT’: Permission denied find: ‘/tmp/nix-3735751-3937767867/scratch-reserved’: Permission denied find: ‘/tmp/nix-3735751-3937767867/u’: Permission denied find: ‘/tmp/nix-3735751-3937767867/tmp’: Permission denied find: ‘/tmp/nix-3735751-3937767867/music’: Permission denied find: ‘/tmp/nix-3735751-3937767867/scratch’: Permission denied find: ‘/tmp/nix-3735751-3937767867/home’: Permission denied find: ‘/tmp/nix-3735751-3937767867/dev’: Permission denied find: ‘/tmp/nix-3735751-3937767867/scratch-hfcs’: Permission denied find: ‘/tmp/nix-3735751-3937767867/srv’: Permission denied find: ‘/tmp/nix-3735751-3937767867/usr’: Permission denied find: ‘/tmp/nix-3735751-3937767867/mnt’: Permission denied find: ‘/tmp/nix-3735751-3937767867/run’: Permission denied find: ‘/tmp/nix-3735751-3937767867/users’: Permission denied find: ‘/tmp/nix-3735751-3937767867/media’: Permission denied find: ‘/tmp/nix-3735751-3937767867/opt’: Permission denied find: ‘/tmp/nix-3735751-3937767867/root’: Permission denied find: ‘/tmp/nix-3735751-3937767867/sys’: Permission denied find: ‘/tmp/nix-3735751-3937767867/var’: Permission denied find: ‘/tmp/nix-3735751-3937767867/lost+found’: Permission denied find: ‘/tmp/nix-3735751-3937767867/proc’: Permission denied find: ‘/tmp/nix-3735751-3937767867/boot’: Permission denied find: ‘/tmp/nix-3735751-3937767867/etc’: Permission denied find: ‘/tmp/nix-shell.c3q1en’: Permission denied find: ‘/tmp/nix-3735094-1866236229/scratch-reserved’: Permission denied find: ‘/tmp/nix-3735094-1866236229/u’: Permission denied find: ‘/tmp/nix-3735094-1866236229/tmp’: Permission denied find: ‘/tmp/nix-3735094-1866236229/music’: Permission denied find: ‘/tmp/nix-3735094-1866236229/scratch’: Permission denied find: ‘/tmp/nix-3735094-1866236229/home’: Permission denied find: ‘/tmp/nix-3735094-1866236229/dev’: Permission denied find: ‘/tmp/nix-3735094-1866236229/scratch-hfcs’: Permission denied find: ‘/tmp/nix-3735094-1866236229/srv’: Permission denied find: ‘/tmp/nix-3735094-1866236229/usr’: Permission denied find: ‘/tmp/nix-3735094-1866236229/mnt’: Permission denied find: ‘/tmp/nix-3735094-1866236229/run’: Permission denied find: ‘/tmp/nix-3735094-1866236229/users’: Permission denied find: ‘/tmp/nix-3735094-1866236229/media’: Permission denied find: ‘/tmp/nix-3735094-1866236229/opt’: Permission denied find: ‘/tmp/nix-3735094-1866236229/root’: Permission denied find: ‘/tmp/nix-3735094-1866236229/sys’: Permission denied find: ‘/tmp/nix-3735094-1866236229/var’: Permission denied find: ‘/tmp/nix-3735094-1866236229/lost+found’: Permission denied find: ‘/tmp/nix-3735094-1866236229/proc’: Permission denied find: ‘/tmp/nix-3735094-1866236229/boot’: Permission denied find: ‘/tmp/nix-3735094-1866236229/etc’: Permission denied find: ‘/tmp/nix-shell.vFcEAv’: Permission denied find: ‘/tmp/nix-3719696-222918794/scratch-reserved’: Permission denied find: ‘/tmp/nix-3719696-222918794/u’: Permission deniedfind: ‘/tmp/nix-3719696-222918794/home’: Permission denied find: ‘/tmp/nix-3719696-222918794/dev’: Permission denied find: ‘/tmp/nix-3719696-222918794/scratch-hfcs’: Permission denied find: ‘/tmp/nix-3719696-222918794/usr’: Permission denied find: ‘/tmp/nix-3719696-222918794/mnt’: Permission denied find: ‘/tmp/nix-3719696-222918794/run’: Permission denied find: ‘/tmp/nix-3719696-222918794/users’: Permission denied find: ‘/tmp/nix-3719696-222918794/media’: Permission denied find: ‘/tmp/nix-3719696-222918794/opt’: Permission denied find: ‘/tmp/nix-3719696-222918794/root’: Permission denied find: ‘/tmp/nix-3719696-222918794/sys’: Permission denied find: ‘/tmp/nix-3719696-222918794/var’: Permission denied find: ‘/tmp/nix-3719696-222918794/lost+found’: Permission denied find: ‘/tmp/nix-3719696-222918794/proc’: Permission denied find: ‘/tmp/nix-3719696-222918794/boot’: Permission denied find: ‘/tmp/nix-3719696-222918794/etc’: Permission denied find: ‘/tmp/nix-shell.eaVKJC’: Permission denied find: ‘/tmp/nix-3719384-1959637618/scratch-reserved’: Permission denied find: ‘/tmp/nix-3719384-1959637618/u’: Permission denied find: ‘/tmp/nix-3719384-1959637618/tmp’: Permission denied find: ‘/tmp/nix-3719384-1959637618/music’: Permission denied find: ‘/tmp/nix-3719384-1959637618/scratch’: Permission denied find: ‘/tmp/nix-3719384-1959637618/home’: Permission denied find: ‘/tmp/nix-3719384-1959637618/dev’: Permission denied find: ‘/tmp/nix-3719384-1959637618/scratch-hfcs’: Permission denied find: ‘/tmp/nix-3719384-1959637618/srv’: Permission denied find: ‘/tmp/nix-3719384-1959637618/usr’: Permission denied find: ‘/tmp/nix-3719384-1959637618/mnt’: Permission denied find: ‘/tmp/nix-3719384-1959637618/run’: Permission denied find: ‘/tmp/nix-3719384-1959637618/users’: Permission denied find: ‘/tmp/nix-3719384-1959637618/media’: Permission denied find: ‘/tmp/nix-3719384-1959637618/opt’: Permission denied find: ‘/tmp/nix-3719384-1959637618/root’: Permission denied find: ‘/tmp/nix-3719384-1959637618/sys’: Permission denied find: ‘/tmp/nix-3719384-1959637618/var’: Permission denied find: ‘/tmp/nix-3719384-1959637618/lost+found’: Permission denied find: ‘/tmp/nix-3719384-1959637618/proc’: Permission denied find: ‘/tmp/nix-3719384-1959637618/boot’: Permission denied find: ‘/tmp/nix-3719384-1959637618/etc’: Permission denied find: ‘/tmp/nix-shell.YOgt3F’: Permission denied find: ‘/tmp/nix-3697404-242710378/scratch-reserved’: Permission denied find: ‘/tmp/nix-3697404-242710378/u’: Permission denied find: ‘/tmp/nix-3697404-242710378/tmp’: Permission denied find: ‘/tmp/nix-3697404-242710378/music’: Permission denied find: ‘/tmp/nix-3697404-242710378/scratch’: Permission denied find: ‘/tmp/nix-3697404-242710378/home’: Permission denied find: ‘/tmp/nix-3697404-242710378/dev’: Permission denied find: ‘/tmp/nix-3697404-242710378/scratch-hfcs’: Permission denied find: ‘/tmp/nix-3697404-242710378/srv’: Permission denied find: ‘/tmp/nix-3697404-242710378/usr’: Permission denied find: ‘/tmp/nix-3697404-242710378/mnt’: Permission denied find: ‘/tmp/nix-3697404-242710378/run’: Permission denied find: ‘/tmp/nix-3697404-242710378/users’: Permission denied find: ‘/tmp/nix-3697404-242710378/media’: Permission denied find: ‘/tmp/nix-3697404-242710378/opt’: Permission denied find: ‘/tmp/nix-3697404-242710378/root’: Permission denied find: ‘/tmp/nix-3697404-242710378/sys’: Permission denied find: ‘/tmp/nix-3697404-242710378/var’: Permission denied find: ‘/tmp/nix-3697404-242710378/lost+found’: Permission denied find: ‘/tmp/nix-3697404-242710378/proc’: Permission denied find: ‘/tmp/nix-3697404-242710378/boot’: Permission denied find: ‘/tmp/nix-3697404-242710378/etc’: Permission denied find: ‘/tmp/nvim.d49chan’: Permission denied find: ‘/tmp/neoforge_installer14200641399993795207’: Permission denied find: ‘/tmp/tmux-26178’: Permission denied find: ‘/tmp/tmux-25857’: Permission denied find: ‘/tmp/neoforge_installer1354699775951996769’: Permission denied find: ‘/tmp/nix-shell.XdGaTM’: Permission denied find: ‘/tmp/nix-3800763-3144653895/scratch-reserved’: Permission denied find: ‘/tmp/nix-3800763-3144653895/u’: Permission denied find: ‘/tmp/nix-3800763-3144653895/tmp’: Permission denied find: ‘/tmp/nix-3800763-3144653895/music’: Permission denied find: ‘/tmp/nix-3800763-3144653895/scratch’: Permission denied find: ‘/tmp/nix-3800763-3144653895/home’: Permission denied find: ‘/tmp/nix-3800763-3144653895/dev’: Permission denied find: ‘/tmp/nix-3800763-3144653895/scratch-hfcs’: Permission denied find: ‘/tmp/nix-3800763-3144653895/srv’: Permission denied find: ‘/tmp/nix-3800763-3144653895/usr’: Permission denied find: ‘/tmp/nix-3800763-3144653895/mnt’: Permission denied find: ‘/tmp/nix-3800763-3144653895/run’: Permission denied find: ‘/tmp/nix-3800763-3144653895/users’: Permission denied find: ‘/tmp/nix-3800763-3144653895/media’: Permission denied find: ‘/tmp/nix-3800763-3144653895/opt’: Permission denied find: ‘/tmp/nix-3800763-3144653895/root’: Permission denied find: ‘/tmp/nix-3800763-3144653895/sys’: Permission denied find: ‘/tmp/nix-3800763-3144653895/var’: Permission denied find: ‘/tmp/nix-3800763-3144653895/lost+found’: Permission denied find: ‘/tmp/nix-3800763-3144653895/proc’: Permission denied find: ‘/tmp/nix-3800763-3144653895/boot’: Permission denied find: ‘/tmp/nix-3800763-3144653895/etc’: Permission denied find: ‘/tmp/neoforge_installer4484965926313671216’: Permission denied find: ‘/tmp/tmux-25113’: Permission denied find: ‘/tmp/tmux-24421’: Permission denied find: ‘/tmp/mcp-Cyx980’: Permission denied find: ‘/tmp/tmux-25306’: Permission denied find: ‘/tmp/claude-26115’: Permission denied find: ‘/tmp/tmux-25523’: Permission denied find: ‘/tmp/tmux-24423’: Permission denied find: ‘/tmp/tmux-26229’: Permission denied find: ‘/tmp/elasticsearch-13792246022977357563’: Permission denied find: ‘/tmp/ztseguin’: Permission denied find: ‘/tmp/systemd-private-4bf5c6bf2f174d558e9778cbfe5fd4f7-nslcd.service-lkdvGB’: Permission denied find: ‘/tmp/systemd-private-4bf5c6bf2f174d558e9778cbfe5fd4f7-lm-sensors.service-ITfiGM’: Permission denied find: ‘/tmp/systemd-private-4bf5c6bf2f174d558e9778cbfe5fd4f7-irqbalance.service-LlkYM4’: Permission denied exec /bin/bash -lc "ls -la ../runtime/venv ../runtime/uv-python ../runtime/uv-cache find ../runtime/venv -maxdepth 3 -printf '%y %p -> %l\\n' | sed -n '1,100p' find ../runtime/uv-python -maxdepth 4 -printf '%y %p -> %l\\n' | sed -n '1,100p' find ../runtime/uv-cache -maxdepth 6 -type f -name 'uv' -o -type f -name 'python' -o -type f -name 'numpy*' | sed -n '1,100p'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 79ms: ../runtime/uv-cache: total 34 drwxrwxr-x 5 22058 22058 8 Aug 24 03:11 . drwxr-xr-x 8 22058 22058 16 Aug 24 04:09 .. -rw-rw-r-- 1 22058 22058 1 Aug 23 23:32 .gitignore -rw-rw-rw- 1 22058 22058 0 Aug 23 23:32 .lock -rw-rw-r-- 1 22058 22058 43 Aug 23 23:32 CACHEDIR.TAG drwxrwxr-x 2 22058 22058 2 Aug 24 03:11 builds-v0 drwxrwxr-x 3 22058 22058 3 Aug 24 03:11 interpreter-v4 drwxrwxr-x 3 22058 22058 5 Aug 24 03:11 sdists-v9 ../runtime/uv-python: total 18 drwxrwxr-x 3 22058 22058 5 Aug 23 23:33 . drwxr-xr-x 8 22058 22058 16 Aug 24 04:09 .. -rw-rw-r-- 1 22058 22058 1 Aug 23 23:33 .gitignore -rw-rw-rw- 1 22058 22058 0 Aug 23 23:33 .lock drwxrwxr-x 2 22058 22058 2 Aug 23 23:33 .temp ../runtime/venv: total 16 drwxrwxr-x 2 22058 22058 2 Aug 23 23:32 . drwxr-xr-x 8 22058 22058 16 Aug 24 04:09 .. d ../runtime/venv -> d ../runtime/uv-python -> d ../runtime/uv-python/.temp -> f ../runtime/uv-python/.lock -> f ../runtime/uv-python/.gitignore -> exec /bin/bash -lc "jq . artifacts/development/improve-additive-native-cubeful-modeling-v1/test-evidence.json jq . artifacts/development/diagnose-native-cubeful-context-error-structure-v1/test-evidence.json sed -n '1,240p' TASK_RESEARCH.md sed -n '1,240p' pyproject.toml command -v nix-shell || true command -v nix || true command -v uvx || true command -v pip || true printenv | sort | rg 'PYTHON|NIX|VIRTUAL|UV|PATH'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 107ms: { "environment": { "blas_threads": 1, "host": "high-fructose-corn-syrup", "nice": 19, "python": "3.11.2", "workers": 1 }, "runs": [ { "command": "pytest -q tests/test_native_cubeful_experiment.py tests/test_position_value_modeling.py", "duration_seconds": 6.75, "passed": 10, "status": "PASS" }, { "command": "pytest -q tests/test_full_corpus_integration.py", "duration_seconds": 0.38, "passed": 2, "skipped": 1, "status": "PASS" } ], "scope": "Experiment-focused unit, boundary, reconstruction, deterministic-fold, and adjacent position-model tests plus the existing full-corpus integration module.", "status": "PASS", "version": "improve-additive-native-cubeful-modeling-v1-test-evidence-v1" } { "commands": [ { "command": "pytest -q tests/test_native_cubeful_context_diagnostic.py tests/test_native_cubeful_experiment.py tests/test_position_value_modeling.py", "failed": 0, "passed": 14, "skipped": 0 }, { "command": "pytest -q tests/test_full_corpus_integration.py", "failed": 0, "passed": 2, "skipped": 1 }, { "command": "git diff --check", "status": "PASS" } ], "identity_sha256": "15c110d98fa598bdbd9ba23511aaccad4694e3244951807a877b7595501181dc", "numerical_thread_limit": 1, "protected_accesses": 0, "status": "PASS", "version": "diagnose-native-cubeful-context-error-structure-v1-test-evidence-v1" } # Research Brain Task ## Full-corpus match-context feasibility **Status:** complete; poor feasibility result **Branch:** `research-brain` **Accepted baseline:** `556793600a579895e61b1209a6500e06bacc249e` **Implementation commit:** `d45b9a7efe04629594574d6651b8aee634666854` ### Scope completed - Parsed all 82 accepted per-game GNU reviews without changing the accepted parser. - Accounted for 3,612 checker headers: 3,258 candidate-bearing decisions and 354 forced cannot-move blocks. - Parsed all 13,929 displayed candidates with zero source failures. - Decoded and round-tripped every GNU Position ID and Match ID. - Legally reconstructed all 13,929 candidate boards with zero quarantines. - Preserved the provisional candidate-set vocabulary and added explicit displayed-count evidence without treating it as legal enumeration. - Built a 50-feature, target-free registry and one canonical group per comparable actual-4-ply decision. - Derived 8,707 unordered comparisons from 2,136 grouped decisions. - Used leave-one-complete-mirrored-pair-out validation across all 10 pairs. - Ran fixed Ridge, interpretable additive, and gradient-boosting experiments. - Simulated a strict explanation-abstention policy. ### Result The feasibility classification is **poor**. The interpretable additive model achieved: - pairwise sign accuracy: 0.5917; - pairwise MAE: 0.0228; - top-move agreement: 0.4813; - top-two ordering accuracy: 0.5646; - mean within-decision rank correlation: 0.2352. The fixed abstention policy supported 174 of 2,136 eligible decisions (8.15%) and abstained on 91.85%. Agreement among supported decisions is 100% by policy construction. Parsing, reconstruction, identifier separation, perspective, and candidate depth preservation passed. The evidence instead points to missing tactical features, insufficient additive capacity, and match-context heterogeneity. ### Decision - [x] Do not call this a money model. - [x] Do not authorize production money generation. - [x] Pause large-scale data generation. - [x] Revise the bounded concepts and narrow the supported explanation scope before considering a genuine money-game pilot. ### Outputs - Contract: `docs/contracts/match-context-feasibility-v1.md` - Builder: `scripts/build_match_context_feasibility.py` - Identifier/reconstruction/features/models: `src/backgammon_explainer/` - Tests: `tests/test_*` - Derived release: `artifacts/derived/sage_gnu_match_context_feasibility_v1/` - Handoff: `docs/handoffs/research/2026-07-22-match-context-feasibility.md` --- ## Existing-data GNU candidate-parser proof **Status:** complete **Branch:** `research-brain` **Implementation commit:** `2279884f8784bd12de10de8a08dcccea5400702b` ## Scope completed - Parsed the two accepted Pair 01 / Match A per-game GNU review fixtures. - Represented all 103 checker-decision blocks. - Represented all 461 displayed candidates. - Preserved each candidate's actual 0-ply, 2-ply, or 4-ply depth. - Kept benchmark `experiment_match_id` separate from GNU `Match ID`. - Added source filenames and one-based line numbers to every parsed row. - Added strict quarantine behavior and focused malformed-input tests. - Produced development-only decision, candidate, validation, and failure files. - Documented a proposed candidate-set vocabulary without freezing it. ## Acceptance result - [x] Existing Sage-error extractor unchanged. - [x] Accepted Stage 1 sources and derived artifacts unchanged. - [x] Every checker decision in the two-file fixture accounted for. - [x] Every displayed candidate represented; no real-source quarantines. - [x] Actual candidate ply preserved. - [x] Mixed-depth decisions explicitly identified. - [x] Parsed rows trace to accepted logical source file and line. - [x] Validation and quarantine outputs are visible. - [x] Nine focused unit tests pass. - [x] Only 2 of 82 review files processed. - [x] No GNU invocation, generation, analysis rerun, or production work. - [x] No off-list move evaluation or model training. ## Pilot findings | Measure | Count | |---|---:| | Checker decisions | 103 | | Displayed candidates | 461 | | Actual 4-ply candidates | 227 | | Actual 2-ply candidates | 177 | | Actual 0-ply candidates | 57 | | Complete full-depth decisions | 1 | | Mixed-depth decisions | 66 | | Uniform incomplete-depth decisions | 18 | | Fewer-than-five displayed decisions | 18 | | Parse failures | 0 | The selected nominal 4-ply review is therefore not a five-candidate 4-ply dataset. Actual per-candidate depth must remain authoritative. ## Outputs - Parser: `src/backgammon_explainer/gnu_review_parser.py` - Builder: `scripts/build_candidate_parser_pilot.py` - Contract: `docs/contracts/gnu-review-candidate-pilot-schema.md` - Fixtures: `tests/fixtures/gnu_review/pair_01_match_A/` - Tests: `tests/test_gnu_review_parser.py` - Pilot: `artifacts/development/candidate_parser_pilot/` ## Next review decision Review the proposed status vocabulary, especially the name `fewer_than_requested_legal_moves`. The pilot observes fewer displayed ranks but does not independently reconstruct the legal-move count. Do not start a larger parse or any GNU money-game work as part of this completed task. [build-system] requires = ["setuptools>=61"] build-backend = "setuptools.build_meta" [project] name = "backgammon-explainer" version = "0.1.0" description = "Auditable backgammon candidate-data research tools" requires-python = ">=3.7" [tool.setuptools] package-dir = {"" = "src"} [tool.setuptools.packages.find] where = ["src"] /users/a2andrad/.local/bin/uvx /users/a2andrad/.local/bin/pip NODE_PATH=/usr/lib/nodejs:/usr/lib/node_modules:/usr/share/javascript PATH=/users/a2andrad/.codex/packages/standalone/releases/0.149.0-x86_64-unknown-linux-musl/codex-path:/users/a2andrad/.local/bin:/users/a2andrad/.local/bin:/usr/local/bin:/usr/bin:/bin:/usr/local/games:/usr/games exec /bin/bash -lc "ls -la /users/a2andrad/.local/bin | sed -n '1,120p' /users/a2andrad/.local/bin/uv --version 2>/dev/null || true /users/a2andrad/.local/bin/uvx --version 2>/dev/null || true /users/a2andrad/.local/bin/pip --version 2>/dev/null || true head -n 5 /users/a2andrad/.local/bin/uvx /users/a2andrad/.local/bin/pip 2>/dev/null || true" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 631ms: total 120982 drwxr-xr-x 2 22058 22058 44 Aug 22 14:38 . drwxr-xr-x 10 22058 22058 10 Aug 15 10:47 .. -rwxr-xr-x 1 22058 22058 188830144 Jul 23 21:52 agy -rwxr-xr-x 1 22058 22058 862248 Jul 10 14:36 bearoffdump lrwxrwxrwx 1 22058 22058 60 Aug 22 14:38 codex -> /users/a2andrad/.codex/packages/standalone/current/bin/codex -rwxr-xr-x 1 22058 22058 219 Jul 9 2020 f2py -rwxr-xr-x 1 22058 22058 219 Jul 9 2020 f2py2 -rwxr-xr-x 1 22058 22058 219 Jul 9 2020 f2py2.7 lrwxrwxrwx 1 22058 22058 41 Jul 16 16:29 fpp -> /users/a2andrad/.local/src/PathPicker/fpp -rwxr-xr-x 1 22058 22058 3821848 Jul 10 14:36 gnubg -rwxr-xr-x 1 22058 22058 233 Oct 27 2015 iptest -rwxr-xr-x 1 22058 22058 233 Oct 27 2015 iptest2 -rwxr-xr-x 1 22058 22058 226 Oct 27 2015 ipython -rwxr-xr-x 1 22058 22058 226 Oct 27 2015 ipython2 lrwxrwxrwx 1 22058 22058 38 Aug 16 16:56 joplin -> /users/a2andrad/.joplin-bin/bin/joplin -rwxr-xr-x 1 22058 22058 215 Oct 27 2015 jsonschema -rwxr-xr-x 1 22058 22058 221 Oct 27 2015 jupyter -rwxr-xr-x 1 22058 22058 220 Oct 27 2015 jupyter-console -rwxr-xr-x 1 22058 22058 263 Oct 27 2015 jupyter-kernelspec -rwxr-xr-x 1 22058 22058 221 Oct 27 2015 jupyter-migrate -rwxr-xr-x 1 22058 22058 98 Oct 28 2015 jupyter-nbconvert -rwxr-xr-x 1 22058 22058 222 Oct 27 2015 jupyter-nbextension -rwxr-xr-x 1 22058 22058 221 Oct 27 2015 jupyter-notebook -rwxr-xr-x 1 22058 22058 223 Oct 27 2015 jupyter-qtconsole -rwxr-xr-x 1 22058 22058 254 Oct 27 2015 jupyter-trust -rwxr-xr-x 1 22058 22058 5468 Nov 25 2015 lolcat -rwxr-xr-x 1 22058 22058 868328 Jul 10 14:36 makebearoff -rwxr-xr-x 1 22058 22058 832872 Jul 10 14:36 makehyper -rwxr-xr-x 1 22058 22058 86960 Jul 10 14:36 makeweights -rwxr-xr-x 1 22058 22058 225 Feb 4 2020 pip -rwxr-xr-x 1 22058 22058 204 Oct 30 2015 pip2 -rwxr-xr-x 1 22058 22058 204 Oct 30 2015 pip2.7 -rwxr-xr-x 1 22058 22058 225 Feb 4 2020 pip3 -rwxr-xr-x 1 22058 22058 225 Feb 4 2020 pip3.5 lrwxrwxrwx 1 22058 22058 45 Jul 26 18:01 quarto -> /users/a2andrad/opt/quarto-1.10.15/bin/quarto lrwxrwxrwx 1 22058 22058 54 Jul 14 16:04 ranger -> /users/a2andrad/.local/pipx/venvs/ranger-fm/bin/ranger lrwxrwxrwx 1 22058 22058 53 Jul 14 16:04 rifle -> /users/a2andrad/.local/pipx/venvs/ranger-fm/bin/rifle -rwxr-xr-x 1 22058 22058 235 Nov 10 2015 tensorboard -rwxr-xr-x 1 22058 22058 245 Nov 10 2015 tensorflow_model_cifar10_eval -rwxr-xr-x 1 22058 22058 256 Nov 10 2015 tensorflow_model_cifar10_multi_gpu_train -rwxr-xr-x 1 22058 22058 246 Nov 10 2015 tensorflow_model_cifar10_train -rwxr-xr-x 1 22058 22058 244 Nov 10 2015 tensorflow_model_mnist_convolutional -rwxr-xr-x 1 22058 22058 65697112 Jul 21 20:58 uv -rwxr-xr-x 1 22058 22058 351120 Jul 21 20:58 uvx uv 0.11.31 (x86_64-unknown-linux-gnu) uvx 0.11.31 (x86_64-unknown-linux-gnu) pip 25.1.1 from /usr/lib/python3/dist-packages/pip (python 3.13) ==> /users/a2andrad/.local/bin/uvx <== ELF>��@PS@8 @! @@@���������������``  ' '��(�#�S�S� j    ��?�?��R�td  ' '��(P�tdx�x�x�� � Q�td���DD/lib64/ld-linux-x86-64.so.2GNU GNU�%5*�Rt�ҳ(�yq�k��ڱ " > X"gns����������������� #0CYnt~��������-; BJQXjx������� ����� &,7  A0iP�PP&y OS&y W``' _ui yui yii �ti �ii �������� �J__libc_start_main__gmon_start___ITM_deregisterTMCloneTable_ITM_registerTMCloneTable__cxa_finalizemallocfree_Unwind_Resumestrlenclosememcpyenvironmemcmpdup2__errno_locationsetgroupssetgidsetuidchrootchdirsetpgidsetsidsignalexecvpreallocpollfcntlopen64dupsysconfpthread_selfpthread_getattr_nppthread_attr_getstackpthread_attr_destroyabortsigactionpauseposix_memalignbcmpmemmovememset_Unwind_RaiseException_Unwind_DeleteException_Unwind_GetLanguageSpecificData_Unwind_GetIPInfo_Unwind_GetRegionStart_Unwind_SetGR_Unwind_SetIPgettidsyscallgetenvgetcwd_Unwind_Backtrace_Unwind_GetIPlseek64readdl_iterate_phdrmunmapmmap64realpathcallocstatxreadlinkwrite_Unwind_GetTextRelBase_Unwind_GetDataRelBasesigaltstackgetauxvalmprotect__xpg_strerror_rpipe2__fxstat64__xstat64libgcc_s.so.1GCC_3.0GCC_3.3GCC_4.2.0libpthread.so.0GLIBC_2.2.5libc.so.6GLIBC_2.3GLIBC_2.3.4GLIBC_2.9GLIBC_2.14GLIBC_2.16 ' '('�w8'�wH'`'�Ph'�'O�'�P�'H'�'`>�'�>�'P>�'`>�'p>�'Vs�'Vs(0M(�m((PM@(pMH(�MP(�NX(i�(b�(�d�(�d�(�o�(�o�(�o�(�o)�o)lr0)lrH)�g`)�gx)k�)�r�)�r�)St�)St�)�i*�i *�b8*�bP*�nh*�n�*�n�*�v�*�v�*�n�*�n�*�n+�n(+�n@+�nX+�np+�n�+�n�+�n�+�n�+�n�+�n,�n,�n0,�nH,�n`,�nx,�n�,�n�,�n�,�n�,�n�,�n-�n -�n8-�nP-�nh-�n�-�n�-�n�-r�-�t�-�t�-�t.�t(.�t@.�tX.sp.s�.��.��.�!/�/�t /�t8/�tP/�th/�t�/�t�/�t�/�t�/�t�/�t�/�t0xk(0xk@0xkX0xkp0xk�0xk�0xk�0xk�0xk�0xk1xk1xk01xkH1xk`1xkx1xk�1xk�1xk�1xk�1xk212�!2�" 2s82�pP2Wvh2n�2���20��2���2���2 e�2 e3�k3d03�kH3 e`3 ex32g�3 e�3�o�3�o�3�o�3��4ee 4p�84�@4�H4��P4p�h4�p4�x4�4p��4P��4���40��4��4�l�4�k5�k5Ԍ05 eH5�t`5�tx5ee�5n�5n�52g�5q�5�t6�t 6��86У@6@�H6P�P6_qh6�t�62g�6p��6���6Х�6��6Е�6С�6��6�7P�7Ki 7KiP7�lX7�bp7}h�7}h�7d��7�l�7�v�7bc8 e8 e08 eH8 e`8 ex8 e�8�h�8�h�8�h�8Lo�8_q9�c 9�c89�cP9�ch9�o�9 e�9 e�9�t�9�t�9���9�e:eu(:eu@: eX:]bp:�c�:�g�:�l�:Ó�:Е�:��:p�;�=;�= ;rP;p�X;��`;���;���;���;���;n��;�e�;Nn�;gd�;�j<�j(<�j@<�jX<�jp<vw�<`��<�k�<�l�<f�<�o=�� =��(=0�@=�PH=0�`=Oh=�Pp=(=x=p��=P��=`��=p��=���=�c�=�c�=�c�=�c>�c>�c0>�cH>�c`>�lx>�o�>ee�>@��>�w�>�w�>c��>=��>`��>P��>K��>[�?E�?U�?x?�w ? �8?B�@?R�H?ޕP?g�`?H�h?X�p?f�x?`�?���?�`APhA�pA�>CpPC`�S���]�"xA�A�A�A�A�A�A �A ==> /users/a2andrad/.local/bin/pip <== #!/usr/bin/python3 # -*- coding: utf-8 -*- import re import sys exec /bin/bash -lc "ls -d /nix/store/*python3*numpy* /nix/store/*python3*pyarrow* /nix/store/*python3*scikit* 2>/dev/null | sed -n '1,80p' ls -d /nix/store/*python3*pytest* 2>/dev/null | sed -n '1,40p' find /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823 -maxdepth 5 -type d -name site-packages -print 2>/dev/null | sed -n '1,80p' /users/a2andrad/.local/bin/uv python list --only-installed 2>/dev/null | sed -n '1,80p'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 377ms: exec /bin/bash -lc "ls /usr/lib/python3/dist-packages | rg 'numpy|pyarrow|sklearn|scipy|pytest|pandas|duckdb' | sed -n '1,120p' find /usr/local/lib -maxdepth 4 -type d \\( -name 'numpy' -o -name 'pyarrow' -o -name 'sklearn' \\) -print 2>/dev/null | sed -n '1,80p' find /opt -maxdepth 5 -type d \\( -name 'numpy' -o -name 'pyarrow' -o -name 'sklearn' \\) -print 2>/dev/null | sed -n '1,80p'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 110ms: exec /bin/bash -lc "/users/a2andrad/.local/bin/uv cache dir ls -la /users/a2andrad/.cache/uv 2>/dev/null | sed -n '1,120p' find /users/a2andrad/.cache/uv -maxdepth 5 -type d \\( -name 'numpy' -o -name 'pyarrow' -o -name 'sklearn' \\) -print 2>/dev/null | sed -n '1,100p' find /users/a2andrad/.cache/uv -maxdepth 4 -type f | rg 'numpy|pyarrow|scikit.learn|scipy|pytest' | sed -n '1,100p'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 239ms: /users/a2andrad/.cache/uv total 51 drwxr-xr-x 9 22058 22058 12 Aug 23 11:43 . drwxrwxr-x 37 22058 22058 40 Aug 19 20:13 .. -rw-r--r-- 1 22058 22058 1 Jul 23 16:54 .gitignore -rw-rw-rw- 1 22058 22058 0 Jul 23 16:54 .lock -rw-r--r-- 1 22058 22058 43 Jul 23 16:54 CACHEDIR.TAG drwxr-xr-x 25 22058 22058 25 Aug 23 11:43 archive-v0 drwxr-xr-x 2 22058 22058 2 Aug 23 11:43 builds-v0 drwxr-xr-x 3 22058 22058 3 Aug 23 11:42 environments-v2 drwxr-xr-x 4 22058 22058 4 Aug 23 10:16 interpreter-v4 drwxr-xr-x 3 22058 22058 5 Jul 23 17:06 sdists-v9 drwxr-xr-x 3 22058 22058 3 Jul 23 16:54 simple-v24 drwxr-xr-x 3 22058 22058 3 Jul 23 16:54 wheels-v6 /users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d/sklearn /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow /users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI/numpy-2.4.6.dist-info/licenses/numpy /users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI/numpy /users/a2andrad/.cache/uv/wheels-v6/pypi/numpy /users/a2andrad/.cache/uv/wheels-v6/pypi/pyarrow /users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d/scikit_learn.libs/libgomp-e985bcbb.so.1.0.0 /users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d/scikit_learn-1.9.0.dist-info/RECORD /users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d/scikit_learn-1.9.0.dist-info/METADATA /users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d/scikit_learn-1.9.0.dist-info/WHEEL /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy-1.17.1.dist-info/WHEEL /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy-1.17.1.dist-info/LICENSE.txt /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy-1.17.1.dist-info/METADATA /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy-1.17.1.dist-info/RECORD /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy/version.py /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy/__config__.py /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy/conftest.py /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy/__init__.py /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy/_distributor_init.py /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy/_cyutility.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy.libs/libquadmath-828275a7.so.0.0.0 /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy.libs/libscipy_openblas-6cdc3b4a.so /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy.libs/libquadmath-96973f99-934c22de.so.0.0.0 /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy.libs/libgfortran-8f1e9814.so.5.0.0 /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4/scipy.libs/libgfortran-040039e1-0352e75f.so.5.0.0 /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow-25.0.1.dist-info/RECORD /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow-25.0.1.dist-info/WHEEL /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow-25.0.1.dist-info/METADATA /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow_substrait.so.2500 /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/error.pxi /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_dataset_orc.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_dataset.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_hdfs.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_dataset_parquet.pxd /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_parquet.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/lib.pxd /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_orc.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_pyarrow_cpp_tests.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/util.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/compat.pxi /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_parquet.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/compute.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow_python_parquet_encryption.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_cuda.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_feather.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_fs.pxd /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_csv.pxd /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_parquet_encryption.pxd /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_substrait.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/json.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/device.pxi /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow.so.2500 /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_hdfs.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libparquet.so.2500 /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_flight.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_dataset_parquet_encryption.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/__init__.pxd /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_dataset_parquet.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow_acero.so.2500 /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/array.pxi /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_compute.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_json.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/types.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/cuda.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_json.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow_python_flight.so.2500 /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/gandiva.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_dlpack.pxi /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_csv.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_orc.pxd /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_parquet_encryption.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/io.pxi /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/scalar.pxi /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow_python.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow_dataset.so.2500 /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/memory.pxi /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_substrait.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_azurefs.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow_flight.so.2500 /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_acero.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow_python.so.2500.1.0 /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow_python.so.2500 /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/substrait.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_compute.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/acero.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_s3fs.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/benchmark.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_dataset_orc.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/pandas_compat.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_generated_version.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_csv.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow_python_flight.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_parquet_encryption.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/lib.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_fs.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_dataset.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_fs.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/csv.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_azurefs.cpython-311-x86_64-linux-gnu.so /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/libarrow_python_parquet_encryption.so.2500.1.0 /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/tensor.pxi /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_compute_docstrings.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_dataset_parquet_encryption.pyx /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/ipc.py /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/public-api.pxi /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y/pyarrow/_dataset.pxd exec /bin/bash -lc "rg -n \"def stable_json|class RegressionMetrics|def result\\(\" src/backgammon_explainer/canonical_analysis.py src/backgammon_explainer/position_value_experiment.py | sed -n '1,120p' sed -n '1,100p' src/backgammon_explainer/canonical_analysis.py sed -n '160,260p' src/backgammon_explainer/position_value_experiment.py sed -n '1,180p' scripts/run_native_cubeful_context_diagnostic.py sed -n '820,1050p' src/backgammon_explainer/native_cubeful_context_diagnostic.py" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 81ms: src/backgammon_explainer/position_value_experiment.py:543:class RegressionMetrics: src/backgammon_explainer/position_value_experiment.py:560: def result(self) -> dict[str, Any]: src/backgammon_explainer/position_value_experiment.py:616: def result(self) -> dict[str, Any]: src/backgammon_explainer/canonical_analysis.py:87:def stable_json(value: Any, *, pretty: bool = False) -> str: """Canonical Analysis Parquet v1 package and interoperability helpers. This module contains the source-neutral behavior shared by commissioning builders and downstream conformance tests. Source ingestion remains outside this module; nullable normalization, context projection, package verification, and immutable promotion do not. """ from __future__ import annotations from dataclasses import dataclass import hashlib import json import math import os import shutil import subprocess import sys import tempfile from pathlib import Path from typing import Any, Callable, Mapping, Sequence import numpy as np import pyarrow as pa import pyarrow.parquet as pq from .gnu_ids import decode_match_id CONTRACT_VERSION = "canonical-analysis-parquet-v1" CONTEXT_ENCODING_VERSION = "occurrence-context-player-on-roll-v1" CANONICAL_TABLE_NAMES = ( "positions", "games", "source_occurrences", "occurrence_contexts", "decisions", "candidates", "evaluations", "cube_occurrences", "cube_actions", "exclusions", ) @dataclass(frozen=True) class CanonicalAnalysisStorageProfile: """Physical Parquet choices for the commissioned Canonical Analysis profile.""" profile_id: str compression: str use_dictionary: bool max_row_group_rows: int parquet_files: tuple[str, ...] def descriptor(self) -> dict[str, Any]: return { "profile_id": self.profile_id, "compression": self.compression, "use_dictionary": self.use_dictionary, "max_row_group_rows": self.max_row_group_rows, "parquet_files": list(self.parquet_files), } CANONICAL_ANALYSIS_STORAGE_PROFILE_V1 = CanonicalAnalysisStorageProfile( profile_id="canonical-analysis-storage-profile-v1", compression="zstd", use_dictionary=True, max_row_group_rows=50_000, parquet_files=tuple(f"{name}.parquet" for name in CANONICAL_TABLE_NAMES), ) FORBIDDEN_CANONICAL_COLUMN_FRAGMENTS = ( "feature_", "_feature", "model_output", "model_prediction", "embedding", "split_assignment", "feature_contribution", "explanation_text", ) def stable_json(value: Any, *, pretty: bool = False) -> str: if pretty: return json.dumps(value, indent=2, sort_keys=True, ensure_ascii=False) + "\n" return json.dumps(value, sort_keys=True, separators=(",", ":"), ensure_ascii=False) def sha256_file(path: Path) -> str: digest = hashlib.sha256() with path.open("rb") as source: for chunk in iter(lambda: source.read(1024 * 1024), b""): digest.update(chunk) return digest.hexdigest() "activity_boundary": {"new_gnu_computations": 0, "new_source_matches": 0, "new_labels": 0, "new_sage_vs_gnu_data": 0}, } payload["identity_sha256"] = _sha256_json(payload) _write_json(output_path, payload) return payload def _sha256_json(value: Any) -> str: return hashlib.sha256(stable_json(value).encode()).hexdigest() def _deterministic_metric_view(value: Any) -> Any: """Normalize insignificant reduction-order noise for rerun identities.""" if isinstance(value, float): return round(value, 12) if isinstance(value, dict): return {key: _deterministic_metric_view(item) for key, item in value.items()} if isinstance(value, list): return [_deterministic_metric_view(item) for item in value] return value def _write_json(path: Path, value: Any) -> None: path.parent.mkdir(parents=True, exist_ok=True) path.write_text(stable_json(value, pretty=True), encoding="utf-8") def _git_head(path: Path) -> str: return subprocess.run( ["git", "rev-parse", "HEAD"], cwd=path, check=True, text=True, stdout=subprocess.PIPE, ).stdout.strip() def _empty_stats(width: int = len(CUBEFUL_REGISTRY), targets: int = len(TARGETS)) -> dict[str, Any]: return { "n": 0, "x_sum": np.zeros(width), "x_sq_sum": np.zeros(width), "xtx": np.zeros((width, width)), "y_sum": np.zeros(targets), "y_sq_sum": np.zeros(targets), "xty": np.zeros((width, targets)), "identity": { "max_probability_identity_error": 0.0, "max_cubeless_identity_error": 0.0, "wrong_perspective_transform_rows": 0, "non_reconstructed_rows": 0, }, } def _merge_stats(destination: dict[str, Any], source: Mapping[str, Any]) -> None: destination["n"] += int(source["n"]) for key in ("x_sum", "x_sq_sum", "xtx", "y_sum", "y_sq_sum", "xty"): destination[key] += source[key] for key in destination["identity"]: if key.startswith("max_"): destination["identity"][key] = max(destination["identity"][key], source["identity"][key]) else: destination["identity"][key] += int(source["identity"][key]) def _update_stats(state: dict[str, Any], x: np.ndarray, y: np.ndarray, *, lose: np.ndarray, transform: Sequence[str], reconstructed: Sequence[str]) -> None: if not len(x): return state["n"] += len(x) state["x_sum"] += x.sum(axis=0) state["x_sq_sum"] += (x * x).sum(axis=0) state["xtx"] += x.T @ x state["y_sum"] += y.sum(axis=0) state["y_sq_sum"] += (y * y).sum(axis=0) state["xty"] += x.T @ y identity = state["identity"] identity["max_probability_identity_error"] = max( identity["max_probability_identity_error"], float(np.max(np.abs(lose - (1.0 - y[:, 0])))) ) derived = y[:, :5] @ PROBABILITY_WEIGHTS - 1.0 identity["max_cubeless_identity_error"] = max( identity["max_cubeless_identity_error"], float(np.max(np.abs(derived - y[:, 5]))) ) identity["wrong_perspective_transform_rows"] += sum( value != "gnu-candidate-to-static-on-roll-v1" for value in transform ) identity["non_reconstructed_rows"] += sum(value != "reconstructed" for value in reconstructed) def _membership(manifest: Mapping[str, Any]) -> tuple[ dict[tuple[str, str], dict[str, int]], dict[tuple[str, str], set[str]], set[str] ]: selection = manifest["selection"] boundaries = [int(selection["checkpoints"][label]["game_prefix_length"]) for label in CHECKPOINTS] train: dict[tuple[str, str], dict[str, int]] = {} for index, item in enumerate(selection["train_game_order"]): campaign, host, worker, game_key = str(item["game_id"]).split("\0", 3) bucket = bisect.bisect_right(boundaries, index) if bucket >= len(CHECKPOINTS): raise RuntimeError("training game lies beyond full checkpoint") train.setdefault((host, worker), {})[game_key] = bucket holdout: dict[tuple[str, str], set[str]] = {} #!/usr/bin/env python3 """Phase runner for diagnose-native-cubeful-context-error-structure-v1.""" from __future__ import annotations import argparse import json from pathlib import Path from backgammon_explainer.native_cubeful_context_diagnostic import ( build_manifest, build_summary, freeze_receipts, record_preflight, score_diagnostic, verify_determinism, verify_package, ) SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/002-diagnose-native-cubeful-context-error-structure-v1.md") SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") RESULT = Path("results/diagnose-native-cubeful-context-error-structure-v1.json") def main() -> int: parser = argparse.ArgumentParser() parser.add_argument("phase", choices=( "freeze", "score", "summarize", "verify-determinism", "manifest", "verify", )) parser.add_argument("--root", type=Path, default=ROOT) args = parser.parse_args() root = args.root if args.phase == "freeze": result = freeze_receipts( protocol_path=PROTOCOL, split_manifest=SPLIT, generation_1_root=GENERATION_1_ROOT, reference_models=REFERENCE_ROOT / "models.json", reference_development=REFERENCE_ROOT / "direct-cubeful.json", output_root=root, ) elif args.phase == "score": record_preflight(evidence_root=root, phase="score-development-diagnostic", storage_path=Path("../runtime")) result = score_diagnostic( shallow_root=SHALLOW, split_manifest=SPLIT, reference_models=REFERENCE_ROOT / "models.json", generation_1_root=GENERATION_1_ROOT, output_root=root, ) elif args.phase == "summarize": result = build_summary(evidence_root=root, result_path=RESULT) elif args.phase == "verify-determinism": result = verify_determinism(root) elif args.phase == "manifest": result = build_manifest(root) else: result = verify_package(root) print(json.dumps(result, indent=2, sort_keys=True)) return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_SCORING" else 1 if __name__ == "__main__": raise SystemExit(main()) "fold_assignment_content_sha256_int8": fold_hash.hexdigest(), "rows": rows, }), } for name, payload in artifacts.items(): _write(output_root / name, payload) run_record = _with_identity({ "version": VERSION + "-run-record-v1", "status": "PASS_COMPLETE_DEVELOPMENT_DIAGNOSTIC", "input_receipt_sha256": sha256_file(receipt_path), "field_mapping_receipt_sha256": sha256_file(mapping_path), "development_candidates": rows, "development_decisions": len(decisions), "development_complete_game_groups": len(group_membership), "elapsed_seconds": time.time() - started, "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, "model_fits": 0, "modeling_candidates": [], "protected_accesses": 0, "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, }) _write(output_root / "run-record.json", run_record) return routing def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: receipt = json.loads((evidence_root / "input-receipt.json").read_text()) routing = json.loads((evidence_root / "routing-record.json").read_text()) stability = json.loads((evidence_root / "support-stability.json").read_text()) access = json.loads((evidence_root / "protected-access-log.json").read_text()) stable = [item for item in stability["stability_cells"] if item["stable_context_regression"]] payload = _with_identity({ "version": VERSION + "-result-v1", "status": "PASS_COMPLETE", "terminal_classification": routing["classification"], "routing_cell": routing["routing_cell"], "stable_context_regression_cells": stable, "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, "development_membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, "prediction_stream_logical_identities": { stream: receipt["prediction_streams"][stream]["logical_prediction_identity_sha256"] for stream in STREAMS }, "model_fits": 0, "modeling_candidates": [], "protected_access_count": len(access["accesses"]), "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", "production": "UNCHANGED", "analyzer": "UNCHANGED", "canonical": "UNCHANGED", "corpus": "UNCHANGED", "candidate_disposition": None, "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", }) _write(evidence_root / "result-summary.json", payload) _write(result_path, payload) return payload def verify_determinism(evidence_root: Path) -> dict[str, Any]: names = ( "field-mapping-receipt.json", "input-receipt.json", "protected-access-log.json", "segment-membership.json", "cell-metrics.json", "fold-evidence.json", "regression-mass.json", "support-stability.json", "routing-record.json", "generation-1-reproduction.json", "prediction-stream-identities.json", "run-record.json", "result-summary.json", "test-evidence.json", ) checks = [] payloads = {} for name in names: path = evidence_root / name value = json.loads(path.read_text()) if path.exists() else {} payloads[name] = value checks.append({"name": "internal_identity:" + name, "status": "PASS" if value and _identity_valid(value) else "FAIL"}) membership = payloads["segment-membership.json"] for family in FAMILY_CELLS: total = sum(item["rows"] for item in membership["families"][family].values()) checks.append({"name": "exhaustive_rows:" + family, "status": "PASS" if total == EXPECTED_DEVELOPMENT_ROWS else "FAIL"}) stability_cells = payloads["support-stability.json"]["stability_cells"] rerouted = route_decision(stability_cells) routing = payloads["routing-record.json"] checks.append({ "name": "deterministic_terminal_routing", "status": "PASS" if all(routing[key] == rerouted[key] for key in rerouted) else "FAIL", }) checks.extend(( { "name": "generation_1_prediction_metric_reproduction", "status": "PASS" if payloads["generation-1-reproduction.json"]["maximum_absolute_metric_error"] <= 1e-12 else "FAIL", }, { "name": "protected_access_exactly_zero", "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", }, { "name": "no_model_fitting", "status": "PASS" if payloads["run-record.json"]["model_fits"] == 0 and not payloads["run-record.json"]["modeling_candidates"] else "FAIL", }, { "name": "accepted_architecture_unchanged", "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", }, { "name": "calculated_cubeful_blocked", "status": "PASS" if payloads["result-summary.json"]["calculated_cubeful"] == "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" else "FAIL", }, )) result = _with_identity({ "version": VERSION + "-deterministic-verification-v1", "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", "checks": checks, "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, "python": platform.python_version(), }) _write(evidence_root / "deterministic-verification.json", result) return result def build_manifest(evidence_root: Path) -> dict[str, Any]: excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) entries = [ {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} for path in files ] payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) payload["package_identity_sha256"] = payload.pop("identity_sha256") _write(evidence_root / "manifest.json", payload) (evidence_root / "SHA256SUMS").write_text( "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", ) return payload def verify_package(evidence_root: Path) -> dict[str, Any]: manifest = json.loads((evidence_root / "manifest.json").read_text()) value = dict(manifest) package_identity = value.pop("package_identity_sha256") checks = [{ "name": "manifest_identity", "status": "PASS" if package_identity == _sha(value) else "FAIL", }] for item in manifest["files"]: path = evidence_root / item["path"] checks.append({ "name": "hash:" + item["path"], "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", }) deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) payload = { "version": VERSION + "-self-verification-v1", "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", "package_identity_sha256": package_identity, "checks": checks, "python": platform.python_version(), } _write(evidence_root / "self-verification.json", payload) return payload exec /bin/bash -lc "sed -n '520,590p' src/backgammon_explainer/position_value_experiment.py sed -n '1,220p' src/backgammon_explainer/native_cubeful_context_diagnostic.py sed -n '220,420p' src/backgammon_explainer/native_cubeful_context_diagnostic.py" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 0ms: "statistics_npz_sha256": sha256_file(statistics_npz), "models": [model.descriptor() for model in models], } payload["identity_sha256"] = _sha256_json(payload) _write_json(output_path, payload) return payload def load_models(path: Path) -> list[RidgeHeads]: payload = json.loads(path.read_text()) models = [] for item in payload["models"]: models.append(RidgeHeads( feature_set=item["feature_set"], checkpoint=item["checkpoint"], feature_ids=tuple(item["feature_ids"]), targets=tuple(item["targets"]), mean=np.asarray(item["mean"]), scale=np.asarray(item["scale"]), standardized_coefficients=np.asarray(item["standardized_coefficients"]), raw_coefficients=np.asarray(item["raw_coefficients"]), intercept=np.asarray(item["intercept"]), alpha=float(item["alpha"]), )) return models class RegressionMetrics: def __init__(self) -> None: self.n = 0 self.sse = self.sae = self.residual_sum = 0.0 self.pred_sum = self.truth_sum = self.pred_sq = self.truth_sq = self.cross = 0.0 def add(self, prediction: np.ndarray, truth: np.ndarray) -> None: prediction, truth = np.asarray(prediction), np.asarray(truth) residual = prediction - truth self.n += len(residual) self.sse += float(residual @ residual) self.sae += float(np.abs(residual).sum()) self.residual_sum += float(residual.sum()) self.pred_sum += float(prediction.sum()); self.truth_sum += float(truth.sum()) self.pred_sq += float(prediction @ prediction); self.truth_sq += float(truth @ truth) self.cross += float(prediction @ truth) def result(self) -> dict[str, Any]: if not self.n: return {"rows": 0} pred_var = self.pred_sq - self.pred_sum ** 2 / self.n truth_var = self.truth_sq - self.truth_sum ** 2 / self.n covariance = self.cross - self.pred_sum * self.truth_sum / self.n correlation = covariance / math.sqrt(pred_var * truth_var) if pred_var > 0 and truth_var > 0 else None return { "rows": self.n, "rmse": math.sqrt(self.sse / self.n), "mae": self.sae / self.n, "bias": self.residual_sum / self.n, "r2": 1.0 - self.sse / truth_var if truth_var > 0 else None, "correlation": correlation, } class PositionModelMetrics: def __init__(self) -> None: self.probability = [RegressionMetrics() for _ in range(5)] self.direct = RegressionMetrics() self.derived = RegressionMetrics() self.agreement = RegressionMetrics() self.outside = np.zeros(5, dtype=np.int64) self.order_violations = {key: 0 for key in ( "win_backgammon_gt_win_gammon_or_better", "win_gammon_or_better_gt_win", "lose_backgammon_gt_lose_gammon_or_worse", "lose_gammon_or_worse_gt_lose", )} self.calibration_count = np.zeros((5, 10), dtype=np.int64) """Frozen Generation 2 diagnostic for native-Cubeful context errors. This module never fits a model and never opens PROTECTED FINAL EVALUATION. It reproduces the three Generation 1 DEVELOPMENT prediction streams and applies only the segment, support, metric, stability, and routing rules frozen in diagnose-native-cubeful-context-error-structure-v1. """ from __future__ import annotations import hashlib import json import math import os import platform import resource import socket import time from dataclasses import dataclass, field from datetime import datetime, timezone from pathlib import Path from typing import Any, Mapping, Sequence import numpy as np import pyarrow.parquet as pq from .canonical_analysis import sha256_file, stable_json from .native_cubeful_experiment import ( DEVELOPMENT_FOLD_SEED, DEVELOPMENT_FOLD_VERSION, EXPECTED_DEVELOPMENT_DECISIONS, EXPECTED_DEVELOPMENT_ROWS, MODEL_IDS, _accepted_baseline, _fold, load_native_models, ) from .position_value_experiment import ( SOURCE_COLUMNS, _candidate_files, _membership, _partition_key, position_classes, ) from .position_value_modeling import cubeful_context_matrix, position_feature_matrix VERSION = "diagnose-native-cubeful-context-error-structure-v1" STARTING_IMPLEMENTATION_HEAD = "2b2a82284649485b00621cd243dfa17c6accca9c" PROTOCOL_FREEZE_COMMIT = "b4c08ba46fbee1f331b5cefe1f787cac09664336" PROTOCOL_SHA256 = "36ff14d5004357deb7d1e57f0b416e408ac93cae2297a600114562fc82a95a69" GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" DEVELOPMENT_COMPLETE_GAMES = 3_178 SUPPORT_MINIMUM_ROWS = 25_000 SUPPORT_MINIMUM_GROUPS = 500 STREAMS = ("accepted_baseline", *MODEL_IDS) PRIMARY_CANDIDATE = "native-cubeful-p3-context-additive-ridge-v1" DIAGNOSTIC_COLUMNS = tuple(dict.fromkeys((*SOURCE_COLUMNS, "candidate_id"))) FAMILY_CELLS = { "cube_ownership": ( "centered", "on_roll_player_owned", "opponent_owned", "unavailable_or_unknown", ), "cube_value": ( "cube_1", "cube_2", "cube_4", "cube_8", "cube_16", "cube_32_plus", "unavailable_or_unknown", ), "match_length": ( "money_or_zero", "length_1_3", "length_4_7", "length_8_15", "length_16_plus", "unavailable_or_unknown", ), "score_relation": ( "tied_away", "on_roll_ahead", "on_roll_behind", "money_or_zero", "unavailable_or_unknown", ), "crawford_state": ( "crawford", "post_crawford", "ordinary_match", "money_or_zero", "unavailable_or_unknown", ), "position_class": ("race", "contact", "bearoff", "other_or_unknown"), } def _sha(value: Any) -> str: return hashlib.sha256(stable_json(value).encode()).hexdigest() def _write(path: Path, value: Any) -> None: path.parent.mkdir(parents=True, exist_ok=True) path.write_text(stable_json(value, pretty=True), encoding="utf-8") def _utc_now() -> str: return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: payload["identity_sha256"] = _sha(payload) return payload def _identity_valid(payload: Mapping[str, Any]) -> bool: value = dict(payload) observed = value.pop("identity_sha256", None) return observed == _sha(value) def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: memory: dict[str, int] = {} for line in Path("/proc/meminfo").read_text().splitlines(): key, raw = line.split(":", 1) memory[key] = int(raw.strip().split()[0]) stat = os.statvfs(storage_path) payload = { "recorded_at_utc": _utc_now(), "phase": phase, "host": socket.gethostname(), "load_average": list(os.getloadavg()), "logical_host_cpus": os.cpu_count(), "process_affinity_cpus": len(os.sched_getaffinity(0)), "mem_total_kib": memory["MemTotal"], "mem_available_kib": memory["MemAvailable"], "storage_path": str(storage_path.resolve()), "storage_free_bytes": stat.f_bavail * stat.f_frsize, "nice": os.getpriority(os.PRIO_PROCESS, 0), "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", "protected_process_actions": [], "numerical_thread_limit": 1, "disposition": "PASS_SUBSTANTIAL_HEADROOM", } evidence_root.mkdir(parents=True, exist_ok=True) with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: stream.write(stable_json(payload) + "\n") return payload def field_mapping_receipt() -> dict[str, Any]: return _with_identity({ "version": VERSION + "-field-mapping-v1", "status": "FROZEN_BEFORE_SCORING", "perspective": "modeled post-move next player on roll", "accepted_factual_decoder": "cubeful_context_matrix(gnu_match_id_native)", "families": { "cube_ownership": { "raw_fields": [ "cubeful_cube_centered", "cubeful_cube_owned_by_player", "cubeful_cube_owned_by_opponent", ], "mapping": { "[1,0,0]": "centered", "[0,1,0]": "on_roll_player_owned", "[0,0,1]": "opponent_owned", "all_other": "unavailable_or_unknown", }, }, "cube_value": { "raw_field": "cubeful_cube_value", "mapping": { "1": "cube_1", "2": "cube_2", "4": "cube_4", "8": "cube_8", "16": "cube_16", ">=32": "cube_32_plus", "all_other": "unavailable_or_unknown", }, }, "match_length": { "raw_fields": ["cubeful_is_money", "cubeful_match_length"], "mapping": { "money_or_length<=0": "money_or_zero", "1..3": "length_1_3", "4..7": "length_4_7", "8..15": "length_8_15", ">=16": "length_16_plus", "all_other": "unavailable_or_unknown", }, }, "score_relation": { "raw_fields": [ "cubeful_match_length", "cubeful_player_score", "cubeful_opponent_score", ], "derived_fields": { "on_roll_points_away": "match_length - player_score", "opponent_points_away": "match_length - opponent_score", }, "mapping": { "equal_points_away": "tied_away", "on_roll_fewer_points_away": "on_roll_ahead", "on_roll_more_points_away": "on_roll_behind", "money_or_length<=0": "money_or_zero", "all_other": "unavailable_or_unknown", }, }, "crawford_state": { "raw_fields": ["cubeful_is_money", "cubeful_match_length", "cubeful_crawford"], "mapping": { "money_or_length<=0": "money_or_zero", "crawford==1": "crawford", "ordinary_non_crawford_match": "ordinary_match", "invalid_or_missing": "unavailable_or_unknown", }, "post_crawford_mapping": "No accepted explicit post-Crawford field exists; non-Crawford match rows map to ordinary_match.", "field_limitation": "POST_CRAWFORD_EXPLICIT_FIELD_UNAVAILABLE", }, "position_class": { "raw_classifier": "accepted deterministic position_classes(position_id)", "mapping": { "race": "race", "contact": "contact", "bearoff": "bearoff", "bar_or_any_other": "other_or_unknown", }, "status": "AVAILABLE", }, }, "outcome_dependent_fields": [], "external_fields": [], }) def _verify_generation_1_package(root: Path) -> dict[str, str]: manifest_path = root / "manifest.json" manifest = json.loads(manifest_path.read_text()) if manifest["package_identity_sha256"] != GENERATION_1_PACKAGE_IDENTITY: raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 package identity") for item in manifest["files"]: if sha256_file(root / item["path"]) != item["sha256"]: raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 package file") return {item["path"]: item["sha256"] for item in manifest["files"]} def freeze_receipts( def freeze_receipts( *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, reference_models: Path, reference_development: Path, output_root: Path, ) -> dict[str, Any]: """Write immutable mappings and input/prediction/partition receipts before scoring.""" if sha256_file(protocol_path) != PROTOCOL_SHA256: raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen protocol") files = _verify_generation_1_package(generation_1_root) split = json.loads(split_manifest.read_text()) authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) models = json.loads((generation_1_root / "models.json").read_text()) development = json.loads((generation_1_root / "development.json").read_text()) result = json.loads((generation_1_root / "result-summary.json").read_text()) if result["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") dev_authority = authorities["source_and_partition_authority"]["development"] holdout = split["selection"]["holdout"] observed = ( dev_authority["membership_sha256"], holdout["game_membership_sha256"], dev_authority["candidates"], dev_authority["decisions"], dev_authority["complete_games"], development["candidates"], development["decisions"], ) expected = ( DEVELOPMENT_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, ) if observed != expected: raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT membership") if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT folds") if sha256_file(split_manifest) != authorities["source_and_partition_authority"]["split_manifest_sha256"]: raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest") if sha256_file(reference_models) != authorities["model_comparison"]["baseline"]["models_artifact_sha256"]: raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline model artifact") if sha256_file(reference_development) != authorities["model_comparison"]["baseline"]["accepted_development_evidence_sha256"]: raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline DEVELOPMENT artifact") output_root.mkdir(parents=True, exist_ok=True) mapping = field_mapping_receipt() _write(output_root / "field-mapping-receipt.json", mapping) model_by_id = {item["model_id"]: item for item in models["models"]} prediction_models = { "accepted_baseline": { "model_id": authorities["model_comparison"]["baseline"]["id"], "model_identity_sha256": authorities["accepted_baseline"]["model_identity_sha256"], }, **{ model_id: {"model_id": model_id, "model_identity_sha256": model_by_id[model_id]["model_identity_sha256"]} for model_id in MODEL_IDS }, } logical_predictions = {} for stream, model in prediction_models.items(): definition = { "stream": stream, "model": model, "development_membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, "target": "native_cubeful_equity_static_next_player", "perspective_transform": "truth=-native_equity; post-move next player on roll", "source_order": "sorted accepted worker partitions, parquet row order, DEVELOPMENT rows only", } logical_predictions[stream] = {**definition, "logical_prediction_identity_sha256": _sha(definition)} receipt = _with_identity({ "version": VERSION + "-input-receipt-v1", "status": "FROZEN_BEFORE_SCORING", "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, "protocol": { "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, "sha256": PROTOCOL_SHA256, }, "generation_1": { "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, "package_file_sha256": files, "development_identity_sha256": development["identity_sha256"], "models_identity_sha256": models["deterministic_identity_sha256"], }, "source": { "authority": authorities["source_and_partition_authority"]["source"], "source_manifest_sha256": authorities["source_and_partition_authority"]["source_manifest_sha256"], "split_manifest_path": str(split_manifest), "split_manifest_sha256": sha256_file(split_manifest), "split_manifest_identity_sha256": split["manifest_identity_sha256"], "reference_models_sha256": sha256_file(reference_models), "reference_development_sha256": sha256_file(reference_development), }, "partitions": { "train": { **authorities["source_and_partition_authority"]["train"], "outcomes_used_in_diagnostic": False, }, "development": dev_authority, "protected_final_evaluation": { "access_budget": 0, "membership_opened": False, "predictions_opened": False, "path_recorded": False, }, }, "grouped_folds": authorities["development_grouped_folds"], "prediction_streams": logical_predictions, "field_mapping_receipt": { "path": "field-mapping-receipt.json", "sha256": sha256_file(output_root / "field-mapping-receipt.json"), "identity_sha256": mapping["identity_sha256"], }, "experiment_limits": { "model_fits": 0, "modeling_candidates": [], "hyperparameter_searches": 0, "feature_searches": 0, "context_feature_additions": 0, "protected_access_budget": 0, }, }) _write(output_root / "input-receipt.json", receipt) access = _with_identity({ "version": VERSION + "-protected-access-log-v1", "access_budget": 0, "accesses": [], "protected_membership_opened": False, "protected_predictions_opened": False, }) _write(output_root / "protected-access-log.json", access) return receipt def segment_labels(context: np.ndarray, raw_classes: np.ndarray) -> dict[str, np.ndarray]: """Apply the exact prospectively frozen raw-to-semantic mappings.""" count = len(context) unknown = "unavailable_or_unknown" ownership = np.full(count, unknown, dtype=object) one_hot = context[:, 8:11] ownership[np.all(one_hot == (1, 0, 0), axis=1)] = "centered" ownership[np.all(one_hot == (0, 1, 0), axis=1)] = "on_roll_player_owned" ownership[np.all(one_hot == (0, 0, 1), axis=1)] = "opponent_owned" cube = context[:, 6] cube_labels = np.full(count, unknown, dtype=object) for value in (1, 2, 4, 8, 16): cube_labels[cube == value] = f"cube_{value}" cube_labels[np.isfinite(cube) & (cube >= 32)] = "cube_32_plus" money, length = context[:, 0] == 1, context[:, 1] money_or_zero = money | (np.isfinite(length) & (length <= 0)) match = (~money_or_zero) & np.isfinite(length) & (length > 0) length_labels = np.full(count, unknown, dtype=object) length_labels[money_or_zero] = "money_or_zero" length_labels[match & (length <= 3)] = "length_1_3" length_labels[match & (length >= 4) & (length <= 7)] = "length_4_7" length_labels[match & (length >= 8) & (length <= 15)] = "length_8_15" length_labels[match & (length >= 16)] = "length_16_plus" player_score, opponent_score = context[:, 2], context[:, 3] score_valid = match & np.isfinite(player_score) & np.isfinite(opponent_score) player_away, opponent_away = length - player_score, length - opponent_score relation = np.full(count, unknown, dtype=object) relation[money_or_zero] = "money_or_zero" relation[score_valid & (player_away == opponent_away)] = "tied_away" relation[score_valid & (player_away < opponent_away)] = "on_roll_ahead" relation[score_valid & (player_away > opponent_away)] = "on_roll_behind" crawford = context[:, 12] crawford_labels = np.full(count, unknown, dtype=object) crawford_labels[money_or_zero] = "money_or_zero" crawford_labels[match & (crawford == 1)] = "crawford" crawford_labels[match & (crawford == 0)] = "ordinary_match" class_labels = np.full(count, "other_or_unknown", dtype=object) for label in ("race", "contact", "bearoff"): class_labels[np.asarray(raw_classes) == label] = label result = { "cube_ownership": ownership, "cube_value": cube_labels, "match_length": length_labels, "score_relation": relation, "crawford_state": crawford_labels, "position_class": class_labels, } for family, labels in result.items(): if set(np.unique(labels)) - set(FAMILY_CELLS[family]): raise AssertionError(f"unfrozen {family} label") return result @dataclass class MetricAccumulator: n: int = 0 sse: float = 0.0 sae: float = 0.0 residual_sum: float = 0.0 pred_sum: float = 0.0 truth_sum: float = 0.0 pred_sq: float = 0.0 truth_sq: float = 0.0 cross: float = 0.0 def add(self, prediction: np.ndarray, truth: np.ndarray) -> None: prediction = np.asarray(prediction, dtype=float) truth = np.asarray(truth, dtype=float) residual = prediction - truth self.n += len(residual) exec /bin/bash -lc "python3 - <<'PY' import hashlib,json fields=[\"cubeful_is_money\",\"cubeful_match_length\",\"cubeful_player_score\",\"cubeful_opponent_score\",\"cubeful_player_away\",\"cubeful_opponent_away\",\"cubeful_cube_value\",\"cubeful_cube_log2\",\"cubeful_cube_centered\",\"cubeful_cube_owned_by_player\",\"cubeful_cube_owned_by_opponent\",\"cubeful_cube_owner_relative_code\",\"cubeful_jacoby\",\"cubeful_cube_offer_pending\"] print(hashlib.sha256(json.dumps(fields,sort_keys=True,separators=(',',':'),ensure_ascii=False).encode()).hexdigest()) PY" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 122ms: 193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac codex The host has substantial headroom (about 100 GiB RAM available, low single-digit load, and 7.2 TiB workspace free), and both worktrees are clean at their frozen heads. I’m now implementing a phase-gated runner that can freeze/commit input authority before any DEVELOPMENT read, fit exactly one candidate from the accepted Generation 1 cache, and has no protected-evaluation phase at all. apply patch patch: completed /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/scripts/run_native_cubeful_crawford_ablation.py /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src/backgammon_explainer/native_cubeful_crawford_ablation.py /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/tests/test_native_cubeful_crawford_ablation.py diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module exec /bin/bash -lc 'python3 -m py_compile src/backgammon_explainer/native_cubeful_crawford_ablation.py scripts/run_native_cubeful_crawford_ablation.py tests/test_native_cubeful_crawford_ablation.py git diff --check git status --short' in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 420ms: ?? scripts/run_native_cubeful_crawford_ablation.py ?? src/backgammon_explainer/native_cubeful_crawford_ablation.py ?? tests/test_native_cubeful_crawford_ablation.py diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module exec /bin/bash -lc "nice -n 19 /users/a2andrad/.local/bin/uv venv --python ../runtime/python311/usr/bin/python3.11 ../runtime/venv nice -n 19 /users/a2andrad/.local/bin/uv pip install --offline --link-mode copy --cache-dir /users/a2andrad/.cache/uv --python ../runtime/venv/bin/python numpy==2.4.6 scipy==1.17.1 scikit-learn==1.9.0 pyarrow==25.0.1 duckdb==1.5.5 pytest==9.0.2 ../runtime/venv/bin/python -c 'import numpy,scipy,sklearn,pyarrow,duckdb,pytest; print(numpy.__version__, scipy.__version__, sklearn.__version__, pyarrow.__version__, duckdb.__version__, pytest.__version__)'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl exited 127 in 171ms: error: Could not acquire lock Caused by: Could not create temporary file Caused by: Read-only file system (os error 30) at path "/users/a2andrad/.cache/uv/.tmphn6gfz" error: Could not acquire lock Caused by: Could not create temporary file Caused by: Read-only file system (os error 30) at path "/users/a2andrad/.cache/uv/.tmpkBS7wj" /bin/bash: line 3: ../runtime/venv/bin/python: No such file or directory diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module exec /bin/bash -lc "du -sh /users/a2andrad/.cache/uv find /users/a2andrad/.cache/uv/archive-v0 -maxdepth 2 -type d -name '*.dist-info' -printf '%h %f\\n' | sort | sed -n '1,120p'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 3111ms: 572M /users/a2andrad/.cache/uv /users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V pluggy-1.6.0.dist-info /users/a2andrad/.cache/uv/archive-v0/3UkR8X-YaELI3BWL sage_gnu_raw_money-0.0.0.dist-info /users/a2andrad/.cache/uv/archive-v0/5cp_68nvRE-Nkkk2 bgsage-1.3.20260723.dist-info /users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG iniconfig-2.3.0.dist-info /users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK packaging-26.2.dist-info /users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn pygments-2.20.0.dist-info /users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4 scipy-1.17.1.dist-info /users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d scikit_learn-1.9.0.dist-info /users/a2andrad/.cache/uv/archive-v0/Jhp421CNutcwZO1Z trove_classifiers-2026.6.1.19.dist-info /users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w pytest-8.4.2.dist-info /users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt threadpoolctl-3.6.0.dist-info /users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI numpy-2.4.6.dist-info /users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo narwhals-2.25.0.dist-info /users/a2andrad/.cache/uv/archive-v0/bUm3Z3yFDY33hYJM sage_gnu_raw_unlimited-0.0.0.dist-info /users/a2andrad/.cache/uv/archive-v0/ejyFUS1BVoSkwog8 hatchling-1.31.0.dist-info /users/a2andrad/.cache/uv/archive-v0/htmgZWvjqOWcLSjx pathspec-1.1.1.dist-info /users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd joblib-1.5.3.dist-info /users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq duckdb-1.5.5.dist-info /users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y pyarrow-25.0.1.dist-info /users/a2andrad/.cache/uv/archive-v0/tFuPWJ2dVsPTH7Rx editables-0.6.dist-info diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module exec /bin/bash -lc "PYTHONHOME=../runtime/python311/usr PYTHONPATH=src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn ../runtime/python311/usr/bin/python3.11 -c 'import numpy,scipy,sklearn,pyarrow,duckdb,pytest; print(numpy.__version__,scipy.__version__,sklearn.__version__,pyarrow.__version__,duckdb.__version__,pytest.__version__)'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl exited 1 in 2086ms: Traceback (most recent call last): File "", line 1, in File "/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d/sklearn/__init__.py", line 70, in from sklearn.base import clone # noqa: E402 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d/sklearn/base.py", line 20, in from sklearn.utils._metadata_requests import _MetadataRequester, _routing_enabled File "/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d/sklearn/utils/__init__.py", line 9, in from sklearn.utils._chunking import gen_batches, gen_even_slices File "/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d/sklearn/utils/_chunking.py", line 11, in from sklearn.utils._param_validation import Interval, validate_params File "/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d/sklearn/utils/_param_validation.py", line 17, in from sklearn.utils.validation import _is_arraylike_not_scalar File "/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d/sklearn/utils/validation.py", line 14, in import narwhals.stable.v2 as nw ModuleNotFoundError: No module named 'narwhals' diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module exec /bin/bash -lc "PYTHONHOME=../runtime/python311/usr PYTHONPATH=src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo ../runtime/python311/usr/bin/python3.11 -c 'import numpy,scipy,sklearn,pyarrow,duckdb,pytest; print(numpy.__version__,scipy.__version__,sklearn.__version__,pyarrow.__version__,duckdb.__version__,pytest.__version__)'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 6494ms: 2.4.6 1.17.1 1.9.0 25.0.1 1.5.5 8.4.2 diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module apply patch patch: completed /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module exec /bin/bash -lc "chmod 755 ../runtime/run-python311 ../runtime/run-python311 -c 'import numpy,pyarrow,pytest; print(numpy.__version__,pyarrow.__version__,pytest.__version__)' OPENBLAS_NUM_THREADS=1 OMP_NUM_THREADS=1 MKL_NUM_THREADS=1 NUMEXPR_NUM_THREADS=1 nice -n 19 ../runtime/run-python311 -m pytest -q tests/test_native_cubeful_crawford_ablation.py tests/test_native_cubeful_experiment.py tests/test_native_cubeful_context_diagnostic.py" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 9257ms: 2.4.6 25.0.1 8.4.2 ............. [100%] 13 passed in 7.25s diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module exec /bin/bash -lc 'OPENBLAS_NUM_THREADS=1 OMP_NUM_THREADS=1 MKL_NUM_THREADS=1 NUMEXPR_NUM_THREADS=1 nice -n 19 ../runtime/run-python311 scripts/run_native_cubeful_crawford_ablation.py freeze' in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 7967ms: { "activity_boundary": { "analyzer_mutation": false, "canonical_mutation": false, "corpus_mutation": false, "new_generic_0ply": 0, "new_gnu": 0, "new_labels": 0, "new_matches": 0, "new_sage": 0, "production_promotion": false, "sage_gnu_campaign_training_rows": 0 }, "context_order_receipt": { "identity_sha256": "7841075b8e5ca0f0ec57338a17e50f072ea7857ea6811fe42ffc3268eb094961", "path": "context-order-receipt.json", "sha256": "a3d073c364bbc642f0feb8da4baa90054bdbc32f48f55117a3cf2e3b785dc214" }, "experiment_limits": { "alternative_family_ablations": 0, "candidate_count": 1, "feature_searches": 0, "hyperparameter_searches": 0, "model_fits": 1, "protected_access_budget": 0 }, "generation_1": { "package_files_sha256": { "contribution-evidence.json": "8bbf2ead035e69391c8c338993541a9ecd321b123c1a58aa575ff8677996dd02", "development.json": "906f25695ad3321b14b8b16da362dd23a7a5ec26c812022bdb2dfa8c1f81ab95", "frozen-authorities.json": "e2b36923b422505c7357eed2d346c5455a0566901587d73bd590efed3f123cbc", "models.json": "582e0eab87dc901c96f611287dec6eae9759995065ac4be633ff69f4ef6575cc", "protected-access-log.json": "8d18d2e10c46d4db96941f511b95717cf58d1186b547696a97c6f1281898538a", "protected-final.json": "a9e429f40b93e9b7ba43359ccc53b959c6d9e14723e9ae687deff71975a52b75", "result-summary.json": "3adbfd00470cfa2ad28c69ac24e9f980bb76fc1fb879639a3d325b41bc0b9468", "test-evidence.json": "dde02f2be47a681f7d8602457e38de43f5cea7e716d70ea79bcc7c7a1c5476ab", "training-cache.json": "dd861f8c2a234947b477db28f429a3ea9bd8a42b379f26df65fc63fa3f714220", "winner-freeze.json": "ee6d5defca55012fd199d4c51c8af6bf88fbf17bb864591d34c790aa5265913c" }, "package_identity_sha256": "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7", "result_identity_sha256": "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" }, "generation_2": { "package_files_sha256": { "cell-metrics.json": "443976e5b9d53ce70d28ff19b894cd7850a72d5565a4ccc486faadc14270cbb8", "deterministic-verification.json": "6422dcaa74b38037d0938c8cbe9b73675125859d6029420725b64e46622496c6", "field-mapping-receipt.json": "b85ffbf1ca86c2c23c02c8199ffd6e5b84287499cd6e68db30f5ca101ecdb166", "fold-evidence.json": "8fc1a51e8d122079c90a0fdb00c4fb2c0355e47a71a0809ca70580641267e1d3", "generation-1-reproduction.json": "4d1869c777a421fea6a1aee9ac5673470b80be513a8abd3e2ab86afabd0f3b70", "input-receipt.json": "54d9bdf9b233eee82e079e9e2f028a840aae06df6206489065df68d32e5e530d", "prediction-stream-identities.json": "5d385dbe1461f41a28bc9652f894d1dc2a94384b785f1245a3519fcb3575837d", "preflight-log.jsonl": "2aca347084fa2f15105629d2c60d16f899c18552fe0bf44b5918949e8a099837", "protected-access-log.json": "98b79e5c39d569d66b92e714035c8627f0b6cbbca76d689d71de2809b9b8a90c", "regression-mass.json": "c4d9f79180d7f0b085b208ab40e2d144710a7e4ad238ed03138a46eda8e41b87", "result-summary.json": "6c77b74d764cfb5231eb0e2e3678911568d5a6b205c833a6fe6ab1ea68dfed6c", "routing-record.json": "f5f818c25036865e626b858e453b1943d86e643ffbf24f809350d369703f3206", "run-record.json": "3e5279842536728291eade3c8c5fcab40f03fee9454f510c38fbce1690ca00c8", "segment-membership.json": "62e95b3ba50455f00990a6e9dcfa90fca718cd6bf0d5ef6739e4484c58508de5", "support-stability.json": "e290d584852e3a9d16040379620f69a803d838d6927fe576d1fa194a2be3d603", "test-evidence.json": "069da60d9939bbf163d50139093b835fc1b237c0df4d354346a847e3e88df881" }, "package_identity_sha256": "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7", "result_identity_sha256": "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb", "routing_cell": { "cell": "money_or_zero", "family": "crawford_state" }, "terminal_classification": "CONTEXT_FAILURE_LOCALIZED" }, "grouped_folds": { "assignment": "SHA256([version,seed,complete_game_id]) modulo four", "folds": 4, "seed": 20260823, "version": "native-cubeful-development-complete-game-fold-4-v1" }, "identity_sha256": "60e130506c021d87e1a0e4dff8fb34f10b8fe8fd3bf4164652d81ad22ef66d77", "model_authority_receipt": { "identity_sha256": "4228943b546e6f03221fdd8bb0fd7e0b5c16be2f2aa87826b9f209466c243488", "path": "model-authority-receipt.json", "sha256": "6dfb6d6ccaf09ea83d29901863862d5999096ae595f2c5b1e01843d635536f97" }, "partitions": { "development": { "candidates": 2094039, "complete_games": 3178, "decisions": 100015, "membership_sha256": "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" }, "protected_final_evaluation": { "access_budget": 0, "membership_opened": false, "path_recorded": false, "predictions_opened": false, "scored": false }, "train": { "candidates": 20981224, "checkpoint": "1000000", "complete_games": 32228, "decisions": 1000002, "membership_sha256": "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" } }, "protocol": { "freeze_commit": "222c2a50418082c65fd57792564c1dc1cc1ad0f6", "path": "../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md", "sha256": "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" }, "source": { "authority": "accepted GNU 0-ply modeling rows; no new engine work", "source_manifest_sha256": "a756567c6c6e316f0bf45527e516127af4317ab8e3b9ea5aa0156ca2dae1c14d", "split_manifest_identity_sha256": "eb0d571182b529588861731d29a49c26880d93f0170646b16e8974d81576ed0f", "split_manifest_path": "artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json", "split_manifest_sha256": "1d125f02d9c5c7340528e134ee6e2815d3ebe612e9df009ae236b726a92019d6", "training_cache_deterministic_identity_sha256": "586b8625b6d00a2c3ef6c7377df103ff2768de61b8997944b5587243bc511a6e", "training_cache_receipt_sha256": "dd861f8c2a234947b477db28f429a3ea9bd8a42b379f26df65fc63fa3f714220" }, "starting_implementation_head": "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a", "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", "target": { "evaluation_mode": "Cubeful", "id": "native_cubeful_equity_static_next_player", "modeled_perspective": "normalized static post-move next player on roll", "source_field": "native_equity", "source_semantics": "GNU native checker-candidate Cubeful equity is maximized by the checker-move player", "transform": "-native_equity" }, "version": "test-crawford-state-context-family-ablation-v1-frozen-authorities-input-receipt-v1" } diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module exec /bin/bash -lc "git status --short git diff --check jq '{status,identity_sha256,partitions,target,experiment_limits}' artifacts/development/test-crawford-state-context-family-ablation-v1/frozen-authorities-input-receipt.json jq '{status,removed_field,retained_context_count,retained_context_order_sha256}' artifacts/development/test-crawford-state-context-family-ablation-v1/context-order-receipt.json jq '{status,candidate_count,candidate_model_id,regularization_grid,optimizer,hyperparameter_searches,feature_searches,alternative_family_ablations}' artifacts/development/test-crawford-state-context-family-ablation-v1/model-authority-receipt.json OPENBLAS_NUM_THREADS=1 OMP_NUM_THREADS=1 MKL_NUM_THREADS=1 NUMEXPR_NUM_THREADS=1 nice -n 19 ../runtime/run-python311 -m pytest -q tests/test_native_cubeful_crawford_ablation.py tests/test_native_cubeful_experiment.py tests/test_native_cubeful_context_diagnostic.py tests/test_position_value_modeling.py tests/test_full_corpus_integration.py" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 8583ms: ?? artifacts/development/test-crawford-state-context-family-ablation-v1/ ?? scripts/run_native_cubeful_crawford_ablation.py ?? src/backgammon_explainer/native_cubeful_crawford_ablation.py ?? tests/test_native_cubeful_crawford_ablation.py { "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", "identity_sha256": "60e130506c021d87e1a0e4dff8fb34f10b8fe8fd3bf4164652d81ad22ef66d77", "partitions": { "development": { "candidates": 2094039, "complete_games": 3178, "decisions": 100015, "membership_sha256": "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" }, "protected_final_evaluation": { "access_budget": 0, "membership_opened": false, "path_recorded": false, "predictions_opened": false, "scored": false }, "train": { "candidates": 20981224, "checkpoint": "1000000", "complete_games": 32228, "decisions": 1000002, "membership_sha256": "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" } }, "target": { "evaluation_mode": "Cubeful", "id": "native_cubeful_equity_static_next_player", "modeled_perspective": "normalized static post-move next player on roll", "source_field": "native_equity", "source_semantics": "GNU native checker-candidate Cubeful equity is maximized by the checker-move player", "transform": "-native_equity" }, "experiment_limits": { "alternative_family_ablations": 0, "candidate_count": 1, "feature_searches": 0, "hyperparameter_searches": 0, "model_fits": 1, "protected_access_budget": 0 } } { "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", "removed_field": "cubeful_crawford", "retained_context_count": 14, "retained_context_order_sha256": "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" } { "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", "candidate_count": 1, "candidate_model_id": "native-cubeful-p3-context-minus-crawford-additive-ridge-v1", "regularization_grid": [ 100.0 ], "optimizer": { "algorithm": "deterministic streaming Adam on Ridge objective", "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, "epochs": 12, "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction" }, "hyperparameter_searches": 0, "feature_searches": 0, "alternative_family_ablations": 0 } ....................s [100%] 20 passed, 1 skipped in 7.18s diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The pre-DEVELOPMENT gate is now durable and validated: the exact TRAIN/DEVELOPMENT identities reproduce, the retained context order is 14 fields with identity `193270…2cac`, and the runner exposes exactly one candidate with alpha 100 and zero searches. The focused and adjacent suite passes (20 passed, 1 skipped). I’m committing this freeze point now, before the fit or any Generation 3 DEVELOPMENT scoring. exec /bin/bash -lc 'git add artifacts/development/test-crawford-state-context-family-ablation-v1/context-order-receipt.json artifacts/development/test-crawford-state-context-family-ablation-v1/frozen-authorities-input-receipt.json artifacts/development/test-crawford-state-context-family-ablation-v1/model-authority-receipt.json artifacts/development/test-crawford-state-context-family-ablation-v1/protected-access-log.json scripts/run_native_cubeful_crawford_ablation.py src/backgammon_explainer/native_cubeful_crawford_ablation.py tests/test_native_cubeful_crawford_ablation.py && git commit -m "Freeze Crawford context ablation implementation" && git status --short --branch && git rev-parse HEAD' in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 242ms: [research/explainer-native-cubeful-k001 6fc2840] Freeze Crawford context ablation implementation 7 files changed, 1686 insertions(+) create mode 100644 artifacts/development/test-crawford-state-context-family-ablation-v1/context-order-receipt.json create mode 100644 artifacts/development/test-crawford-state-context-family-ablation-v1/frozen-authorities-input-receipt.json create mode 100644 artifacts/development/test-crawford-state-context-family-ablation-v1/model-authority-receipt.json create mode 100644 artifacts/development/test-crawford-state-context-family-ablation-v1/protected-access-log.json create mode 100644 scripts/run_native_cubeful_crawford_ablation.py create mode 100644 src/backgammon_explainer/native_cubeful_crawford_ablation.py create mode 100644 tests/test_native_cubeful_crawford_ablation.py ## research/explainer-native-cubeful-k001...origin/research/explainer-native-cubeful-k001 [ahead 1] 6fc28407dadfd93c72f048eae3418cedc3d976aa diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The authority freeze is committed at `6fc2840`, satisfying the required commit-before-DEVELOPMENT boundary. I’m starting the single permitted fit now from the accepted 5.9 GiB Generation 1 cache, at nice 19 with one numerical thread; this phase has no DEVELOPMENT or protected-data reader. exec /bin/bash -lc 'OPENBLAS_NUM_THREADS=1 OMP_NUM_THREADS=1 MKL_NUM_THREADS=1 NUMEXPR_NUM_THREADS=1 nice -n 19 ../runtime/run-python311 scripts/run_native_cubeful_crawford_ablation.py fit' in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module succeeded in 707199ms: /users/a2andrad/.local/lib/python2.7/site-packages/numpy /users/a2andrad/bg-env/lib/python3.13/site-packages/numpy /users/a2andrad/bg-env/lib/python3.11/site-packages/numpy /users/a2andrad/workspaces/explainer-k001/.venv/lib/python3.11/site-packages/numpy /users/a2andrad/sage-gnu-raw-unlimited-runner/.venv/lib/python3.11/site-packages/numpy /users/a2andrad/.cache/uv/archive-v0/zt2jrZvxRcumjutg/lib/python3.11/site-packages/numpy /users/a2andrad/.cache/uv/archive-v0/XQMM7ZgZ-JadMC8A/lib/python3.11/site-packages/numpy /users/a2andrad/backgammon-explainer-gnu0ply-modeling/.venv-modeling/lib/python3.11/site-packages/numpy /users/a2andrad/scratch/backgammon-explainer-deep-4ply-acquisition-v1/.venv/lib/python3.11/site-packages/numpy /users/a2andrad/venv/gwem/lib/python2.7/site-packages/numpy /users/a2andrad/venv/data-science/lib/python2.7/site-packages/numpy /users/a2andrad/venv/trading/trading/lib/python2.7/site-packages/numpy /users/a2andrad/canonical-explainer-conversion/.venv/lib/python3.11/site-packages/numpy /users/a2andrad/gnu-gnu-raw-unlimited-runner/.venv/lib/python3.11/site-packages/numpy /users/a2andrad/.local/bin/uv uv 0.11.31 (x86_64-unknown-linux-gnu) diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit is still running cleanly with no output or error, which is expected because the frozen trainer writes only after all 12 epochs complete. I’m leaving the process undisturbed and continuing short polls so Corpus/Post Match retain priority. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single fit remains active and silent after roughly two minutes; there is still no indication of failure or resource contention. No other process has been touched, and DEVELOPMENT remains unopened for Generation 3. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The trainer continues normally. Because it is deterministic and checkpoint-free within this frozen run, I’m not interrupting it to inspect intermediate epochs; the durable model and full epoch record will be written atomically at completion. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex At about four minutes, the fit remains active with no errors. The process is constrained to the requested single BLAS thread and nice 19, so the elapsed time is consistent with yielding CPU priority on HFCS. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The run remains stable and deliberately low-priority. I’ll continue monitoring it without changing scheduling or touching any priority workload; the next substantive checkpoint is the completed candidate artifact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Six minutes in, there is still no error output and the fit has not produced a partial model. This preserves the frozen one-fit protocol: I’m waiting for the complete 12-epoch artifact rather than adapting mid-run. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The trainer remains healthy and silent at roughly seven minutes. For scale, this is processing 20,981,224 accepted TRAIN candidates across 12 deterministic passes; no new labels or engine work are involved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit continues without deviation. I’m maintaining the exact frozen optimizer schedule and avoiding any parallel model work, search, or diagnostic that could alter the protocol or consume extra HFCS headroom. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex At around nine minutes, the single candidate fit is still running cleanly at low priority. The Generation 3 protected-access log remains at exactly zero, and scoring has not begun. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Ten minutes in, execution remains stable with no warnings or failures. I’m continuing the frozen run unchanged; once it completes I’ll validate the 8,220 optimizer updates and model identity before opening DEVELOPMENT. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The trainer is still active and silent after about eleven minutes, as expected for the existing implementation’s end-only reporting. No protocol boundary has been crossed and no adaptive action has been taken. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-model fit continues normally at roughly twelve minutes. I’m preserving the required substantial headroom and will not accelerate it by increasing threads or priority. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex At about thirteen minutes, the deterministic fit is still progressing without output or error. This remains within the expected cost of 12 passes over the full accepted TRAIN population on one low-priority thread. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains stable around fourteen minutes in. I’m continuing to monitor only this lane’s process; no GNU, Sage, Corpus, Canonical, Analyzer, or production state is being modified. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Fifteen minutes in, the process is still active with no error stream. I’m keeping the running protocol fixed; the next action remains validation of the completed candidate, not any mid-run adaptation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The low-priority fit continues normally after about sixteen minutes. All boundaries remain intact: one candidate, one alpha, no search, and zero protected accesses. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex At roughly seventeen minutes, execution is still clean and unchanged. I’m deliberately accepting the longer runtime imposed by nice 19 and one-thread arithmetic to preserve HFCS priority and reproducibility. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit is still active at about eighteen minutes with no warnings. No partial result has been used, so DEVELOPMENT/ADAPTIVE SELECTION remains uncontaminated by intermediate outcomes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Around nineteen minutes in, the deterministic trainer continues without error. I’m maintaining the exact frozen course and monitoring for completion only. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains active and stable at approximately twenty minutes. The runtime is conservative by design; I’m not increasing concurrency or priority, and no other lane process is being affected. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex At about twenty-one minutes, the single fit is still cleanly running. I’m preserving the frozen optimizer state and will not interrupt or restart it. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The low-priority trainer remains active after roughly twenty-two minutes, still without errors or partial artifacts. All evaluation partitions remain separated as frozen. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex At approximately twenty-three minutes, the one permitted model fit continues normally. I’m keeping monitoring lightweight and leaving the shared host’s priority workloads untouched. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit is still stable near twenty-four minutes. No errors, adaptive changes, extra candidates, or protected accesses have occurred. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex At about twenty-five minutes, the deterministic training phase remains in progress and clean. I’ll continue it unchanged until the candidate model and epoch record are durable. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains active and error-free at roughly twenty-six minutes. The conservative single-thread/nice-19 policy is still in force, with no attempt to borrow priority from Corpus or Post Match. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Around twenty-seven minutes in, the run continues normally and silently. I’m holding the protocol fixed and waiting for completion. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single candidate fit remains clean at about twenty-eight minutes. No evaluation metrics have been inspected, so there has been no within-protocol adaptation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex At approximately twenty-nine minutes, the trainer continues as expected under the conservative host policy. I’m leaving it uninterrupted. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The run has reached roughly thirty minutes and remains active without error. The frozen fit is computationally substantial but bounded; I’m continuing to prioritize reproducibility and HFCS safety over speed. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex At about thirty-one minutes, the fit is still progressing cleanly. All frozen data and access boundaries remain unchanged. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic run remains healthy at roughly thirty-two minutes, still with no error stream or partial result. I’m maintaining the exact one-thread, nice-19 execution. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex At approximately thirty-three minutes, training continues normally and without intervention. No additional model, ablation, or search has been introduced. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single candidate fit remains active after about thirty-four minutes, with all protocol and host-safety constraints intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex At roughly thirty-five minutes, the fit continues cleanly and unchanged. I’m still waiting for the complete durable model before any DEVELOPMENT access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The run remains stable at about thirty-six minutes. I’m checking only the previously recorded Generation 1 elapsed-time metadata—not live optimizer state—to calibrate expectations while leaving the current fit untouched. exec /bin/bash -lc "jq '{elapsed_seconds,peak_rss_kib,updates:.models[1].optimizer.updates,epochs:.models[1].optimizer.epochs}' artifacts/development/improve-additive-native-cubeful-modeling-v1/models.json uptime free -h | sed -n '1,3p'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 111ms: { "elapsed_seconds": 7867.863333940506, "peak_rss_kib": 31280520, "updates": 8220, "epochs": 12 } 05:07:42 up 42 days, 7:51, 15 users, load average: 4.96, 5.31, 5.42 total used free shared buff/cache available Mem: 157Gi 57Gi 17Gi 1.8Gi 86Gi 99Gi Swap: 0B 0B 0B diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Generation 1’s comparable two-model fit took about 131 minutes; the current one-model ablation is therefore behaving within precedent. HFCS still shows about 99 GiB available RAM and modest load, so the safety disposition remains PASS and I’m continuing unchanged. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains active and error-free. Its duration is consistent with the accepted Generation 1 implementation, and no resource-safety condition has changed. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues under the frozen single-thread schedule, with no errors and no partial evaluation access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The candidate fit remains stable at low priority. All one-fit, no-search, and zero-protected-access constraints remain in force. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit continues normally and remains the only active experiment work for this lane. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Execution remains clean and unchanged; I’m continuing to wait for the fully durable 12-epoch candidate artifact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single-model trainer remains active with no warnings. Host headroom and protocol boundaries are unchanged. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The frozen fit is still progressing normally. I’m not changing threads, priority, data, or optimizer settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority, with DEVELOPMENT still unopened for Generation 3 and protected access still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted candidate fit continues without error or intervention. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains active and clean. I’m continuing the frozen protocol unchanged. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit is still proceeding normally at nice 19 with one numerical thread; no safety or protocol condition has changed. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean and unchanged, with no additional candidate or search activity. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single frozen fit continues normally. DEVELOPMENT and PROTECTED FINAL EVALUATION remain untouched. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The low-priority trainer remains stable, and the running protocol is unchanged. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit continues cleanly at the frozen resource limits, with no evidence of failure or host pressure. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate training phase remains active and unchanged. I’m continuing to wait for its complete durable output. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable with no warnings, adaptation, or partition access beyond accepted TRAIN. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues at low priority. All frozen experiment constraints remain satisfied. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The candidate fit remains active and error-free; no other experimental work has been started. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The run continues normally under the single-thread, nice-19 policy, with protected access still at zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The frozen training phase remains stable and unchanged; I’m continuing to monitor for completion. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single candidate continues fitting cleanly; no result-dependent action has been taken. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains healthy around the one-hour mark, still within the precedent set by the 131-minute Generation 1 two-model run. Host and protocol constraints remain intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training is still active and error-free at the frozen low-priority settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single-model fit continues normally with no boundary changes or additional work. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains clean and low-priority. DEVELOPMENT has not yet been opened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit remains in progress without warnings or deviation from the frozen optimizer. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly at the conservative resource limits; no adaptive or external action has occurred. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single candidate fit remains active and stable, with the access log still at zero for PROTECTED FINAL EVALUATION. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The frozen run continues normally, unchanged and low-priority. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The candidate fit is still progressing cleanly; the experiment remains exactly one fit with no search. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains healthy and unchanged under the frozen host policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-model fit continues without error or intervention, with all partition boundaries preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains stable and low-priority; no additional data or model work has been introduced. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit continues cleanly at the frozen settings. DEVELOPMENT remains unopened and protected access remains zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single fit remains active and error-free, with no protocol changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally under the low-priority resource cap. I’m leaving the optimizer uninterrupted. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The frozen one-candidate run remains healthy, with no evaluation or external-system access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The trainer continues cleanly at nice 19 with one numerical thread; all frozen boundaries remain intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains active without warnings, search, or adaptation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single deterministic fit continues normally, still isolated to accepted TRAIN data. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged; protected access is still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly under the frozen resource policy, with no additional work launched. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit remains active and error-free, with all data partitions and authority roles preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally with no warnings or protocol deviation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The low-priority single fit remains stable; DEVELOPMENT and protected data are still untouched. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The candidate fit continues cleanly and unchanged at the frozen host limits. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-model trainer remains active with no errors, searches, or access-boundary changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally at nice 19 with one numerical thread; no other system state has been altered. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit remains stable and within precedent; all frozen constraints remain satisfied. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single candidate continues fitting cleanly, with no DEVELOPMENT or protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains error-free and unchanged under the frozen one-thread schedule. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The frozen fit remains healthy around ninety minutes, still below the Generation 1 two-model duration and with ample host headroom. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single fit continues normally, unchanged and low-priority. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable with all one-fit, no-search, and zero-protected-access rules intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues cleanly, with no partial outcomes inspected and no protocol adaptation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit remains active and error-free at the frozen low-priority settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally and unchanged; protected access remains exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single deterministic fit remains stable with no resource or protocol changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The trainer continues cleanly under the exact frozen schedule; no evaluation boundary has been crossed. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains error-free and unchanged; this is still the sole Generation 3 fit. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The low-priority fit continues normally with no warnings, search, or protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The candidate fit remains stable at about 100 minutes, still consistent with the existing Generation 1 implementation’s cost. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly under the frozen resource limits; no boundary or host-safety condition has changed. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single frozen fit remains active and error-free, with no model or data changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally at low priority. DEVELOPMENT and protected evaluation remain untouched. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit remains clean and unchanged, with all frozen constraints still satisfied. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues without errors or intervention at the frozen low-priority settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit remains active and stable; protected access is still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-model run continues normally with no search, adaptation, or evaluation access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged under the frozen host-safety policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The frozen fit continues cleanly; no protocol boundary has changed. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single deterministic trainer remains active and error-free at the exact frozen settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally with no additional candidates, data, or system changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains stable and low-priority, with DEVELOPMENT and protected evaluation still untouched. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit continues cleanly and unchanged under the frozen protocol. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free; all architecture and access boundaries are unchanged. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally at low priority, with no within-protocol adaptation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single-model fit remains stable and unchanged; protected access remains exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly at the frozen resource settings, with no additional work or access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate run remains healthy and unchanged under the frozen protocol. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues normally, still isolated to accepted TRAIN and still at zero protected accesses. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains stable around two hours, still consistent with the existing implementation’s end-only reporting and low-priority execution. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The runtime is now close to Generation 1’s 131-minute precedent; basis construction dominates, so removing the second model does not halve elapsed time. The process remains clean and within all frozen limits. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally and unchanged, with no errors or access-boundary changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit remains active and clean at the exact frozen resource settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single candidate continues fitting without warning or deviation; no evaluation data has been opened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged, with all frozen experiment constraints intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The low-priority one-model fit continues cleanly; protected access remains exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains active and error-free, with no adaptation or extra candidate work. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally at the frozen nice/thread limits; all data boundaries remain preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic one-candidate fit remains clean and unchanged near the Generation 1 elapsed-time benchmark. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority at about 130 minutes; no warnings or protocol changes have occurred. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit has reached the prior Generation 1 duration and remains clean. I’m continuing without interruption until the end-only artifact is written. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic run continues normally just beyond the earlier two-model elapsed time, still with no error or resource-safety signal. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged; all frozen protocol boundaries are still intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly at low priority; no adaptive action has been taken. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free, with DEVELOPMENT and protected evaluation still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues normally and unchanged at the frozen host limits. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-model trainer remains stable with no warning, search, or access-boundary change. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly, and the frozen one-fit protocol remains unchanged. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The low-priority fit remains active and error-free, with no impact on other lane or host authorities. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains clean at roughly 140 minutes. I’m not treating elapsed time alone as a blocker because host headroom remains substantial and the process has shown no failure signal. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally under the frozen low-priority schedule, with no errors or protocol changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single candidate fit remains stable and unchanged; all protected and production boundaries remain intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly with no adaptive action or extra model work. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains active and error-free at the frozen resource limits. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-model fit continues normally, with DEVELOPMENT and protected evaluation still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable, low-priority, and unchanged under the frozen protocol. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains error-free; the modest overrun versus Generation 1 is acceptable under changing shared-host load and nice-19 scheduling, so I’m continuing without intervention. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally with all access, architecture, and activity boundaries preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single candidate fit remains clean and unchanged; no result-dependent action has occurred. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains active and stable at the frozen low-priority settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly with no warning, adaptation, or evaluation access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit remains stable and low-priority; all frozen boundaries remain intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues without errors or protocol deviation, with protected access still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic one-model fit remains clean and unchanged under the host-safety limits. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally; no other model, family, or data source has been introduced. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains active and error-free at nice 19 with one numerical thread; evaluation partitions remain unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic run continues cleanly. Its extra column-selection copy makes this ablation somewhat slower than the full-context basis path, but changing that implementation mid-protocol would violate the freeze, so I’m continuing unchanged. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged, with no access-boundary or host-safety issue. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues normally and without intervention; protected access remains zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean under the exact frozen optimizer and resource schedule. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues normally, with no adaptive action or extra experiment activity. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single-model trainer remains stable and low-priority; all frozen data and authority boundaries remain intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly with no warnings, searches, or protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit remains active and error-free under the frozen host policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally and unchanged; evaluation remains unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority, with protected access still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues without error or protocol deviation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean at the frozen nice/thread settings, with all boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues normally, with no additional candidates, searches, or access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single model remains in training, clean and low-priority, with protected access still zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues unchanged under the frozen optimizer and host-safety policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains active and error-free; no protocol boundary has changed. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean and low-priority, with all data and architecture boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues normally; no adaptive or protected access has occurred. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable under the exact frozen resource schedule, with no errors or external mutations. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit remains active and clean; DEVELOPMENT and protected evaluation are still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally and unchanged, with the one-fit/no-search boundary intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single candidate fit remains stable and low-priority, with all frozen boundaries preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly with no warning, adaptation, or access-boundary change. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The frozen fit remains healthy around three hours, still with substantial host headroom and no failure signal; I’m continuing unchanged. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally at the exact frozen settings; no result has been exposed or acted upon. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains active and error-free, with all protocol boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly and unchanged under the low-priority host policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable with no errors, searches, or protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit remains active and unchanged, with evaluation partitions still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly at the frozen nice/thread settings; no protocol condition has changed. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit remains stable and error-free, with zero protected access and no other model work. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally and unchanged under the frozen host-safety policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains active and clean, with all frozen access and activity boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The single fit continues without error or deviation; no DEVELOPMENT metric has been inspected. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority; all frozen constraints remain satisfied. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly at the frozen resource limits, with protected access still zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally and unchanged, with no search, adaptation, or external mutation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit remains active and error-free under the frozen host policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority, with all frozen data and authority boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit continues without warnings or protocol changes; evaluation remains unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly at the exact frozen resource settings, with protected access still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains active and unchanged, with no additional model or data work. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean and low-priority, with all frozen protocol boundaries preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains active around 200 minutes with no failure signal. Because the protocol forbids mid-run adaptation and host safety remains acceptable, I’m continuing the exact run rather than restarting or optimizing it. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly under the frozen low-priority settings; no evaluation or external mutation has occurred. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic one-candidate fit remains stable and unchanged, with protected access still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally at nice 19 with one numerical thread; all frozen boundaries remain intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains active and error-free, with no adaptation, search, or extra data consumption. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority; evaluation partitions remain unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues cleanly with all one-fit and zero-protected-access rules intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally under the exact frozen resource and optimizer settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-model fit remains stable and unchanged, with no warning or protocol deviation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly at low priority; protected access remains exactly zero and no external state has changed. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic one-candidate fit remains active and error-free under the frozen host policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally and unchanged, with evaluation still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains stable and low-priority, with all frozen activity and access boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly, with no adaptation, extra candidate, or protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains active and unchanged under the exact frozen settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean and low-priority; all frozen partition and authority boundaries remain preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues normally, with no warning or protocol change. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable under the frozen one-thread schedule, with protected access still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues cleanly, unchanged and isolated to accepted TRAIN data. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free, with no result-dependent action or access-boundary change. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains healthy around 220 minutes. The slower path is still bounded and host-safe, so I’m preserving the frozen run rather than altering its implementation mid-experiment. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally under the exact one-thread/nice-19 schedule, with all boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit remains stable and error-free; no evaluation or external mutation has occurred. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean and unchanged under the frozen protocol, with protected access still zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally, with no search, adaptation, or extra data consumption. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and low-priority, with all frozen access and authority boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-model fit continues cleanly at the exact frozen resource settings; evaluation remains unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged, with no warning or protocol deviation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues normally under the frozen host policy, with protected access still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean and low-priority, with all one-fit/no-search boundaries preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate trainer continues without errors or access-boundary changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged under the exact frozen resource schedule. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues cleanly, with no adaptive or external action and zero protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free; evaluation partitions are still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-model fit continues normally at nice 19 with one numerical thread, with all frozen constraints intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged, with no warning, search, or protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues cleanly under the exact frozen resource and optimizer policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and low-priority; all frozen data, access, and authority boundaries remain intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues normally with no adaptive action or extra model work. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and error-free under the frozen host policy; protected access remains zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains healthy around four hours. This is long but still an active, bounded, host-safe run; I’m preserving the frozen protocol and continuing unchanged. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly. I’m checking only its read-only process state and current host headroom to confirm it is computing rather than stalled; I will not signal or alter it. exec /bin/bash -lc "ps -eo pid,ppid,ni,stat,etime,time,%cpu,%mem,rss,comm,args --sort=-%cpu | sed -n '1,30p' uptime free -h | sed -n '1,3p'" in /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl succeeded in 49ms: PID PPID NI STAT ELAPSED TIME %CPU %MEM RSS COMMAND COMMAND 1 0 19 SNs 00:00 00:00:00 55.5 0.0 16980 codex codex-linux-sandbox --sandbox-policy-cwd /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl --command-cwd /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl --permission-profile {"type":"managed","file_system":{"type":"restricted","entries":[{"path":{"type":"special","value":{"kind":"root"}},"access":"read"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl"},"access":"write"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm"},"access":"write"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.git"},"access":"write"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.git"},"access":"write"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001"},"access":"write"},{"path":{"type":"special","value":{"kind":"slash_tmp"}},"access":"write"},{"path":{"type":"special","value":{"kind":"tmpdir"}},"access":"write"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.git"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.git"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.git/.git"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.git/.git"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/.git"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.agents"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.agents"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.git/.agents"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.git/.agents"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/.agents"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.codex"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.codex"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/.git/.codex"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/tm/.git/.codex"},"access":"read","missing_path_behavior":"skip"},{"path":{"type":"path","path":"/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/.codex"},"access":"read","missing_path_behavior":"skip"}]},"network":"restricted"} --apply-seccomp-then-exec -- /bin/bash -c __CODEX_SNAPSHOT_OVERRIDE_SET_0="${CODEX_APPLY_PATCH_PRESERVE_LINE_ENDINGS+x}" __CODEX_SNAPSHOT_OVERRIDE_0="${CODEX_APPLY_PATCH_PRESERVE_LINE_ENDINGS-}" __CODEX_SNAPSHOT_OVERRIDE_SET_1="${CODEX_PERMISSION_PROFILE+x}" __CODEX_SNAPSHOT_OVERRIDE_1="${CODEX_PERMISSION_PROFILE-}" __CODEX_SNAPSHOT_OVERRIDE_SET_2="${CODEX_PLUGIN_METRICS_OUTPUT+x}" __CODEX_SNAPSHOT_OVERRIDE_2="${CODEX_PLUGIN_METRICS_OUTPUT-}" __CODEX_SNAPSHOT_OVERRIDE_SET_3="${CODEX_SESSION_ID+x}" __CODEX_SNAPSHOT_OVERRIDE_3="${CODEX_SESSION_ID-}" __CODEX_SNAPSHOT_OVERRIDE_SET_4="${CODEX_THREAD_ID+x}" __CODEX_SNAPSHOT_OVERRIDE_4="${CODEX_THREAD_ID-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_0="${ALL_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_0="${ALL_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_1="${BUNDLE_HTTPS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_1="${BUNDLE_HTTPS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_2="${BUNDLE_HTTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_2="${BUNDLE_HTTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_3="${BUNDLE_NO_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_3="${BUNDLE_NO_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_4="${BUNDLE_SSL_CA_CERT+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_4="${BUNDLE_SSL_CA_CERT-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_5="${CARGO_HTTP_CAINFO+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_5="${CARGO_HTTP_CAINFO-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_6="${CODEX_CA_CERTIFICATE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_6="${CODEX_CA_CERTIFICATE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_7="${CODEX_NETWORK_ALLOW_LOCAL_BINDING+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_7="${CODEX_NETWORK_ALLOW_LOCAL_BINDING-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_8="${CODEX_NETWORK_PROXY_ACTIVE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_8="${CODEX_NETWORK_PROXY_ACTIVE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_9="${CODEX_NETWORK_PROXY_ATTRIBUTION+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_9="${CODEX_NETWORK_PROXY_ATTRIBUTION-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_10="${CODEX_NETWORK_PROXY_BROKERED_CREDENTIALS+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_10="${CODEX_NETWORK_PROXY_BROKERED_CREDENTIALS-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_11="${CODEX_NETWORK_PROXY_CREDENTIAL_BROKER_ACTIVE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_11="${CODEX_NETWORK_PROXY_CREDENTIAL_BROKER_ACTIVE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_12="${CURL_CA_BUNDLE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_12="${CURL_CA_BUNDLE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_13="${DOCKER_HTTPS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_13="${DOCKER_HTTPS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_14="${DOCKER_HTTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_14="${DOCKER_HTTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_15="${ELECTRON_GET_USE_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_15="${ELECTRON_GET_USE_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_16="${FTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_16="${FTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_17="${GIT_SSL_CAINFO+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_17="${GIT_SSL_CAINFO-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_18="${HTTPS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_18="${HTTPS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_19="${HTTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_19="${HTTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_20="${NODE_EXTRA_CA_CERTS+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_20="${NODE_EXTRA_CA_CERTS-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_21="${NODE_USE_ENV_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_21="${NODE_USE_ENV_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_22="${NO_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_22="${NO_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_23="${NPM_CONFIG_CAFILE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_23="${NPM_CONFIG_CAFILE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_24="${NPM_CONFIG_HTTPS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_24="${NPM_CONFIG_HTTPS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_25="${NPM_CONFIG_HTTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_25="${NPM_CONFIG_HTTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_26="${NPM_CONFIG_NOPROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_26="${NPM_CONFIG_NOPROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_27="${NPM_CONFIG_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_27="${NPM_CONFIG_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_28="${PIP_CERT+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_28="${PIP_CERT-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_29="${PIP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_29="${PIP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_30="${REQUESTS_CA_BUNDLE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_30="${REQUESTS_CA_BUNDLE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_31="${SSL_CERT_FILE+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_31="${SSL_CERT_FILE-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_32="${WSS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_32="${WSS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_33="${WS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_33="${WS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_34="${YARN_HTTPS_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_34="${YARN_HTTPS_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_35="${YARN_HTTP_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_35="${YARN_HTTP_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_36="${YARN_NO_PROXY+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_36="${YARN_NO_PROXY-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_37="${all_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_37="${all_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_38="${ftp_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_38="${ftp_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_39="${http_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_39="${http_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_40="${https_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_40="${https_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_41="${no_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_41="${no_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_42="${npm_config_cafile+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_42="${npm_config_cafile-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_43="${npm_config_http_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_43="${npm_config_http_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_44="${npm_config_https_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_44="${npm_config_https_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_45="${npm_config_noproxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_45="${npm_config_noproxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_46="${npm_config_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_46="${npm_config_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_47="${ws_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_47="${ws_proxy-}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_48="${wss_proxy+x}" __CODEX_SNAPSHOT_PROXY_OVERRIDE_48="${wss_proxy-}" __CODEX_SNAPSHOT_PROXY_ENV_SET="${CODEX_NETWORK_PROXY_ACTIVE+x}" if . '/users/a2andrad/.codex/shell_snapshots/01a032d1-af70-7903-99d2-14de6b7620fe.1787558997909339012.sh' >/dev/null 2>&1; then :; fi if [ -n "${__CODEX_SNAPSHOT_OVERRIDE_SET_0}" ]; then export CODEX_APPLY_PATCH_PRESERVE_LINE_ENDINGS="${__CODEX_SNAPSHOT_OVERRIDE_0}"; else unset CODEX_APPLY_PATCH_PRESERVE_LINE_ENDINGS; fi if [ -n "${__CODEX_SNAPSHOT_OVERRIDE_SET_1}" ]; then export CODEX_PERMISSION_PROFILE="${__CODEX_SNAPSHOT_OVERRIDE_1}"; else unset CODEX_PERMISSION_PROFILE; fi if [ -n "${__CODEX_SNAPSHOT_OVERRIDE_SET_2}" ]; then export CODEX_PLUGIN_METRICS_OUTPUT="${__CODEX_SNAPSHOT_OVERRIDE_2}"; else unset CODEX_PLUGIN_METRICS_OUTPUT; fi if [ -n "${__CODEX_SNAPSHOT_OVERRIDE_SET_3}" ]; then export CODEX_SESSION_ID="${__CODEX_SNAPSHOT_OVERRIDE_3}"; else unset CODEX_SESSION_ID; fi if [ -n "${__CODEX_SNAPSHOT_OVERRIDE_SET_4}" ]; then export CODEX_THREAD_ID="${__CODEX_SNAPSHOT_OVERRIDE_4}"; else unset CODEX_THREAD_ID; fi if [ -n "$__CODEX_SNAPSHOT_PROXY_ENV_SET" ] || [ -n "${CODEX_NETWORK_PROXY_ACTIVE+x}" ]; then if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_0}" ]; then export ALL_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_0}"; else unset ALL_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_1}" ]; then export BUNDLE_HTTPS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_1}"; else unset BUNDLE_HTTPS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_2}" ]; then export BUNDLE_HTTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_2}"; else unset BUNDLE_HTTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_3}" ]; then export BUNDLE_NO_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_3}"; else unset BUNDLE_NO_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_4}" ]; then export BUNDLE_SSL_CA_CERT="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_4}"; else unset BUNDLE_SSL_CA_CERT; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_5}" ]; then export CARGO_HTTP_CAINFO="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_5}"; else unset CARGO_HTTP_CAINFO; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_6}" ]; then export CODEX_CA_CERTIFICATE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_6}"; else unset CODEX_CA_CERTIFICATE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_7}" ]; then export CODEX_NETWORK_ALLOW_LOCAL_BINDING="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_7}"; else unset CODEX_NETWORK_ALLOW_LOCAL_BINDING; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_8}" ]; then export CODEX_NETWORK_PROXY_ACTIVE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_8}"; else unset CODEX_NETWORK_PROXY_ACTIVE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_9}" ]; then export CODEX_NETWORK_PROXY_ATTRIBUTION="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_9}"; else unset CODEX_NETWORK_PROXY_ATTRIBUTION; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_10}" ]; then export CODEX_NETWORK_PROXY_BROKERED_CREDENTIALS="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_10}"; else unset CODEX_NETWORK_PROXY_BROKERED_CREDENTIALS; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_11}" ]; then export CODEX_NETWORK_PROXY_CREDENTIAL_BROKER_ACTIVE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_11}"; else unset CODEX_NETWORK_PROXY_CREDENTIAL_BROKER_ACTIVE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_12}" ]; then export CURL_CA_BUNDLE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_12}"; else unset CURL_CA_BUNDLE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_13}" ]; then export DOCKER_HTTPS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_13}"; else unset DOCKER_HTTPS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_14}" ]; then export DOCKER_HTTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_14}"; else unset DOCKER_HTTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_15}" ]; then export ELECTRON_GET_USE_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_15}"; else unset ELECTRON_GET_USE_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_16}" ]; then export FTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_16}"; else unset FTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_17}" ]; then export GIT_SSL_CAINFO="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_17}"; else unset GIT_SSL_CAINFO; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_18}" ]; then export HTTPS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_18}"; else unset HTTPS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_19}" ]; then export HTTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_19}"; else unset HTTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_20}" ]; then export NODE_EXTRA_CA_CERTS="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_20}"; else unset NODE_EXTRA_CA_CERTS; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_21}" ]; then export NODE_USE_ENV_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_21}"; else unset NODE_USE_ENV_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_22}" ]; then export NO_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_22}"; else unset NO_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_23}" ]; then export NPM_CONFIG_CAFILE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_23}"; else unset NPM_CONFIG_CAFILE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_24}" ]; then export NPM_CONFIG_HTTPS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_24}"; else unset NPM_CONFIG_HTTPS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_25}" ]; then export NPM_CONFIG_HTTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_25}"; else unset NPM_CONFIG_HTTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_26}" ]; then export NPM_CONFIG_NOPROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_26}"; else unset NPM_CONFIG_NOPROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_27}" ]; then export NPM_CONFIG_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_27}"; else unset NPM_CONFIG_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_28}" ]; then export PIP_CERT="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_28}"; else unset PIP_CERT; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_29}" ]; then export PIP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_29}"; else unset PIP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_30}" ]; then export REQUESTS_CA_BUNDLE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_30}"; else unset REQUESTS_CA_BUNDLE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_31}" ]; then export SSL_CERT_FILE="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_31}"; else unset SSL_CERT_FILE; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_32}" ]; then export WSS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_32}"; else unset WSS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_33}" ]; then export WS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_33}"; else unset WS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_34}" ]; then export YARN_HTTPS_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_34}"; else unset YARN_HTTPS_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_35}" ]; then export YARN_HTTP_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_35}"; else unset YARN_HTTP_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_36}" ]; then export YARN_NO_PROXY="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_36}"; else unset YARN_NO_PROXY; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_37}" ]; then export all_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_37}"; else unset all_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_38}" ]; then export ftp_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_38}"; else unset ftp_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_39}" ]; then export http_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_39}"; else unset http_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_40}" ]; then export https_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_40}"; else unset https_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_41}" ]; then export no_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_41}"; else unset no_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_42}" ]; then export npm_config_cafile="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_42}"; else unset npm_config_cafile; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_43}" ]; then export npm_config_http_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_43}"; else unset npm_config_http_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_44}" ]; then export npm_config_https_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_44}"; else unset npm_config_https_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_45}" ]; then export npm_config_noproxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_45}"; else unset npm_config_noproxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_46}" ]; then export npm_config_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_46}"; else unset npm_config_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_47}" ]; then export ws_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_47}"; else unset ws_proxy; fi if [ -n "${__CODEX_SNAPSHOT_PROXY_OVERRIDE_SET_48}" ]; then export wss_proxy="${__CODEX_SNAPSHOT_PROXY_OVERRIDE_48}"; else unset wss_proxy; fi fi if [ -n "${PATH:-}" ]; then export PATH='/users/a2andrad/.codex/packages/standalone/releases/0.149.0-x86_64-unknown-linux-musl/codex-path':"$PATH"; else export PATH='/users/a2andrad/.codex/packages/standalone/releases/0.149.0-x86_64-unknown-linux-musl/codex-path'; fi exec '/bin/bash' -c 'ps -eo pid,ppid,ni,stat,etime,time,%cpu,%mem,rss,comm,args --sort=-%cpu | sed -n '"'"'1,30p'"'"' uptime free -h | sed -n '"'"'1,3p'"'"'' 2 1 19 SN 00:00 00:00:00 0.0 0.0 3212 bash /bin/bash -c ps -eo pid,ppid,ni,stat,etime,time,%cpu,%mem,rss,comm,args --sort=-%cpu | sed -n '1,30p' uptime free -h | sed -n '1,3p' 3 2 19 RN 00:00 00:00:00 0.0 0.0 3916 ps ps -eo pid,ppid,ni,stat,etime,time,%cpu,%mem,rss,comm,args --sort=-%cpu 4 2 19 SN 00:00 00:00:00 0.0 0.0 2180 sed sed -n 1,30p 09:10:07 up 42 days, 11:53, 15 users, load average: 5.33, 5.71, 5.48 total used free shared buff/cache available Mem: 157Gi 57Gi 16Gi 1.8Gi 87Gi 99Gi Swap: 0B 0B 0B diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The read-only check confirms host headroom is still substantial—about 99 GiB available RAM, no swap pressure, and modest load. The isolated PID namespace cannot expose the fit itself, but the running session is responsive and has no error output, so the safety disposition remains PASS. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly under the frozen settings, with all data and access boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit remains active and error-free, with no protocol change or adaptive action. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority; protected access is still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally under the exact frozen host and optimizer policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly, with no search, adaptation, evaluation access, or external mutation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit remains stable and unchanged, with all frozen boundaries preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free at the frozen low-priority settings; evaluation remains unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic one-model fit continues cleanly, with zero protected access and no protocol changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable under the exact frozen resource schedule, with no adaptive or external action. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit continues normally and unchanged, with all partition and authority boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean and low-priority; no search, adaptation, or evaluation access has occurred. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues without errors or protocol deviation, with protected access still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and stable under the exact frozen host and optimizer settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly and unchanged, with evaluation partitions still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority, with all frozen boundaries preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally, with no additional candidate, search, or protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains clean near twice the Generation 1 elapsed time, consistent with the added retained-column copy. I’m continuing unchanged through completion, as required by the frozen protocol. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally and unchanged, with all access and host-safety boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains active and error-free, with protected access still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable under the exact frozen schedule, with no adaptive or external action. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly, with evaluation partitions still unopened and all frozen boundaries preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free under the frozen low-priority policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues normally, with no search, adaptation, or protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged, with all frozen data and authority boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit continues cleanly at nice 19 with one numerical thread; protected access remains zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free, with no protocol deviation or extra work. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally under the exact frozen resource settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean and unchanged, with evaluation partitions still unopened and protected access at zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit remains stable and low-priority, with all frozen constraints preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally with no error, adaptation, or additional candidate activity. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit remains active and error-free under the exact frozen host policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged, with no evaluation or external-system access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit continues cleanly at nice 19 with one numerical thread; all frozen boundaries remain intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and low-priority, with protected access still exactly zero and no protocol changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally, unchanged and isolated to accepted TRAIN data. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable with no warnings, search, or access-boundary change. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly under the exact frozen resource settings, with evaluation still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free, with all frozen activity and authority boundaries preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally at low priority, with protected access still zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged, with no adaptive action or external mutation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit continues cleanly under the frozen host and optimizer policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free; evaluation partitions remain unopened and protected access remains zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally, with no search, adaptation, or additional data consumption. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable at the exact frozen low-priority settings, with all authority boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly, unchanged and isolated to accepted TRAIN data. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free, with protected access still exactly zero and no protocol changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues normally under the frozen one-thread/nice-19 policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged, with no adaptive action, search, or evaluation access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate trainer continues cleanly at the exact frozen resource settings; protected access remains zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free under the frozen protocol, with all boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues normally, with no additional candidate, data, or external-system activity. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged under the exact frozen host policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit continues cleanly, with evaluation partitions still unopened and protected access at zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free with no protocol deviation or additional work. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally under the exact frozen resource and optimizer settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority, with all frozen data, access, and authority boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly, with no search, adaptation, or protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free at the frozen low-priority settings; evaluation is still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains healthy around five hours. It is still an active, bounded run with ample host headroom and no failure signal, so I’m continuing unchanged to honor the frozen protocol. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly under the exact frozen settings, with no protocol or access-boundary changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate trainer remains stable and error-free; protected access is still exactly zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean and unchanged at nice 19 with one numerical thread, with all frozen boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit continues normally, with no adaptation, search, or evaluation access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and low-priority, with all frozen data and authority boundaries preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit continues cleanly under the exact frozen host and optimizer settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and error-free, with protected access still zero and evaluation still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally, with no extra model, data, or external-system activity. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and unchanged under the frozen low-priority policy, with all boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly, with no search, adaptation, or protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and error-free at the exact frozen resource settings; evaluation remains unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally, with all frozen access and activity boundaries preserved. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean and low-priority, with no protocol change or external mutation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit continues cleanly under the exact frozen optimizer and host settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and error-free; protected access remains exactly zero and evaluation is still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally, with no search, adaptation, or additional data activity. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains clean and low-priority, with all frozen data, access, and authority boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly under the exact frozen resource policy, with no protocol changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and error-free, with evaluation partitions still unopened and protected access at zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The fit remains an active, host-safe process despite the long elapsed time. Because the current implementation and schedule were committed before evaluation, I’m preserving them exactly rather than making a mid-experiment performance change. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues cleanly under the frozen settings, with no errors, adaptation, or evaluation access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate trainer remains stable and low-priority, with all frozen access and authority boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally under the exact frozen host and optimizer policy; protected access remains zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic fit remains active and error-free, with no additional candidate or data activity. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged at the frozen low-priority settings, with evaluation still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly, with no search, adaptation, or protected access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free under the exact frozen resource settings, with all boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally, with no external mutation or evaluation access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority, with protected access still exactly zero and no protocol changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit remains clean around five and a half hours. Host safety is still acceptable, so I’m maintaining the committed protocol and waiting for the durable end result. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training continues normally under the exact frozen settings, with no errors or access-boundary changes. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer remains active and error-free, with all one-fit and zero-protected-access boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged at nice 19 with one numerical thread; evaluation remains unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly, with no search, adaptation, or external mutation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free under the exact frozen resource policy; protected access remains zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally, unchanged and isolated to accepted TRAIN data. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority, with all frozen boundaries intact and no evaluation access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit continues cleanly, with no protocol deviation or additional work. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free under the exact frozen host and optimizer settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally, with protected access still exactly zero and evaluation still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged, with no adaptive action, search, or external mutation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly under the frozen low-priority policy, with all boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free; no evaluation access or additional model work has occurred. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally at nice 19 with one numerical thread; protected access remains zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged under the exact frozen host and optimizer policy. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit continues cleanly, with all frozen data, access, and authority boundaries intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free, with no search, adaptation, or evaluation access. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally under the exact frozen low-priority settings. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and unchanged, with protected access still exactly zero and no external mutation. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one-candidate fit continues cleanly, unchanged and isolated to accepted TRAIN data. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains active and error-free under the frozen host policy; evaluation is still unopened. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The deterministic trainer continues normally, with all one-fit, no-search, and zero-protected-access rules intact. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex Training remains stable and low-priority, with no protocol change or additional work. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module codex The one permitted fit continues cleanly at the exact frozen resource settings; protected access remains zero. diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module diff --git a//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 new file mode 100644 index 0000000000000000000000000000000000000000..bb9b288042bf7e30aeb4b4d16e5d51d3610901a6 --- /dev/null +++ b//users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/run-python311 @@ -0,0 +1,4 @@ +#!/bin/sh +export PYTHONHOME=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr +export PYTHONPATH=/users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/impl/src:/users/a2andrad/.cache/uv/archive-v0/RYRoBHbmfZS40tQI:/users/a2andrad/.cache/uv/archive-v0/qYLNVds7OwQuj61y:/users/a2andrad/.cache/uv/archive-v0/H_TvSJ4EzPM8L7A4:/users/a2andrad/.cache/uv/archive-v0/IwQ5dl34rYH-Q7-d:/users/a2andrad/.cache/uv/archive-v0/nqGdsEyN-7ROEykd:/users/a2andrad/.cache/uv/archive-v0/Q9__Q9Hix9te-7kt:/users/a2andrad/.cache/uv/archive-v0/onOF2j9qL3c3-9tq:/users/a2andrad/.cache/uv/archive-v0/PbSPsJ6HKyhheE8w:/users/a2andrad/.cache/uv/archive-v0/-QeApw68zHzhKe4V:/users/a2andrad/.cache/uv/archive-v0/8fxvQwGFJZue7jpG:/users/a2andrad/.cache/uv/archive-v0/AqlL8rzpFWAykYwK:/users/a2andrad/.cache/uv/archive-v0/G3l95_VOr-BWaSDn:/users/a2andrad/.cache/uv/archive-v0/WK9PB9PvYTJiRMmo +exec /users/a2andrad/scratch/bgs-explainer-24h-v3-20260823/lanes/explainer-native-cubeful-k001/runtime/python311/usr/bin/python3.11 "$@" diff --git a/scripts/run_native_cubeful_crawford_ablation.py b/scripts/run_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3abbda0ed44c214ae514c2691a4b1e1b3f6f7757 --- /dev/null +++ b/scripts/run_native_cubeful_crawford_ablation.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Phase runner for test-crawford-state-context-family-ablation-v1.""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + build_manifest, + build_summary, + fit_candidate, + freeze_receipts, + record_preflight, + score_development, + verify_artifact_package, + verify_determinism, +) + + +SHALLOW = Path("/users/a2andrad/backgammon-explainer-gnu0ply-modeling/artifacts/development/gnu_0ply_modeling_full_1h/run_001") +PROTOCOL = Path("../tm/milestones/explainer-native-cubeful-k001/prompts/003-test-crawford-state-context-family-ablation-v1.md") +SPLIT = Path("artifacts/development/explainer-k002-shallow-to-deep/protocol-v1/split-checkpoint-manifest.json") +REFERENCE_ROOT = Path("artifacts/development/explainer-k002-position-value-modeling") +GENERATION_1_ROOT = Path("artifacts/development/improve-additive-native-cubeful-modeling-v1") +GENERATION_2_ROOT = Path("artifacts/development/diagnose-native-cubeful-context-error-structure-v1") +ROOT = Path("artifacts/development/test-crawford-state-context-family-ablation-v1") +CACHE = Path("../runtime/native-cubeful-training-cache") +RESULT = Path("results/test-crawford-state-context-family-ablation-v1.json") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("phase", choices=( + "freeze", "fit", "score-development", "summarize", + "verify-determinism", "manifest", "verify", + )) + parser.add_argument("--root", type=Path, default=ROOT) + parser.add_argument("--cache-root", type=Path, default=CACHE) + args = parser.parse_args() + root = args.root + if args.phase == "freeze": + result = freeze_receipts( + protocol_path=PROTOCOL, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, generation_2_root=GENERATION_2_ROOT, + generation_1_result=Path("results/improve-additive-native-cubeful-modeling-v1.json"), + generation_2_result=Path("results/diagnose-native-cubeful-context-error-structure-v1.json"), + reference_models=REFERENCE_ROOT / "models.json", + reference_development=REFERENCE_ROOT / "direct-cubeful.json", + output_root=root, + ) + elif args.phase == "fit": + record_preflight(evidence_root=root, phase="fit-single-candidate", storage_path=args.cache_root) + result = fit_candidate( + cache_root=args.cache_root, generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + output_path=root / "candidate-model.json", + ) + elif args.phase == "score-development": + record_preflight(evidence_root=root, phase="score-development", storage_path=Path("../runtime")) + result = score_development( + shallow_root=SHALLOW, split_manifest=SPLIT, + generation_1_root=GENERATION_1_ROOT, + receipt_path=root / "frozen-authorities-input-receipt.json", + model_path=root / "candidate-model.json", output_root=root, + ) + elif args.phase == "summarize": + result = build_summary(evidence_root=root, result_path=RESULT) + elif args.phase == "verify-determinism": + result = verify_determinism(root) + elif args.phase == "manifest": + result = build_manifest(root) + else: + result = verify_artifact_package(root) + print(json.dumps(result, indent=2, sort_keys=True)) + return 0 if result.get("status", "PASS").startswith("PASS") or result.get("status") == "FROZEN_BEFORE_FIT_AND_DEVELOPMENT" else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/backgammon_explainer/native_cubeful_crawford_ablation.py b/src/backgammon_explainer/native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..3a4bf875b8bb0c4513a659071f5ff4edce3d7e50 --- /dev/null +++ b/src/backgammon_explainer/native_cubeful_crawford_ablation.py @@ -0,0 +1,932 @@ +"""Frozen Generation 3 Crawford-context ablation for native Cubeful equity. + +The module fits exactly one candidate from the accepted Generation 1 TRAIN +cache. It removes ``cubeful_crawford`` before additive basis construction, +scores DEVELOPMENT once, and contains no PROTECTED FINAL EVALUATION reader. +""" + +from __future__ import annotations + +import hashlib +import json +import math +import os +import platform +import resource +import socket +import time +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Mapping, Sequence + +import numpy as np +import pyarrow.parquet as pq + +from .canonical_analysis import sha256_file, stable_json +from .constrained_additive_position_model import _adam_step +from .native_cubeful_experiment import ( + ALPHA, + BATCH_SIZE, + DEVELOPMENT_FOLD_SEED, + DEVELOPMENT_FOLD_VERSION, + EPOCHS, + EXPECTED_DEVELOPMENT_DECISIONS, + EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_TRAIN_ROWS, + FULL_WIDTH, + P3_WIDTH, + TARGET, + TRAIN_CHECKPOINT, + _fold, + _iter_cache, + load_cache, +) +from .position_value_experiment import ( + RegressionMetrics, + SOURCE_COLUMNS, + _candidate_files, + _membership, + _partition_key, +) +from .position_value_modeling import ( + CUBEFUL_CONTEXT_REGISTRY, + P3_REGISTRY, + cubeful_context_matrix, + position_feature_matrix, +) + + +VERSION = "test-crawford-state-context-family-ablation-v1" +STARTING_IMPLEMENTATION_HEAD = "75ce3ba6cf9a03db98886b26f1d2dc6b057f6e2a" +PROTOCOL_FREEZE_COMMIT = "222c2a50418082c65fd57792564c1dc1cc1ad0f6" +PROTOCOL_SHA256 = "740c6f3aadb5603e333b8570a3ef37b1a814a12b81da1127fb67e7dcc1f339f4" +GENERATION_1_RESULT_IDENTITY = "6afcf567f77765c8cb428a761024373d19fb80d5b40f1e350e234b99f1d9e8be" +GENERATION_1_PACKAGE_IDENTITY = "65435fd555e701672eba3984404b350084d1b51df7ae5bdaefd3acc2c3a976e7" +GENERATION_2_RESULT_IDENTITY = "b40afe6802c375fcc89d3a34a3cbeef4cec7708576c9c3dce971e44ca11dbddb" +GENERATION_2_PACKAGE_IDENTITY = "d4be9a1562a02b62d4716bcc9cd94bcc3422531f0d483666a81dcc57872143d7" +TRAIN_MEMBERSHIP_IDENTITY = "2a0fe6336e8aa3af5df118bb57f7644de71c51af7200f7ff9977bd5c1541adec" +DEVELOPMENT_MEMBERSHIP_IDENTITY = "5b83e7d3c68666a0b8b15c51bf51450960b3dac5bdb1b6bd65e2599ddc5ab544" +DEVELOPMENT_COMPLETE_GAMES = 3_178 +POSITION_REGISTRY_IDENTITY = "30ede35745bbbc645683f93473ef67cd9e21ff1f152a7369c26498c350fd0287" +FULL_CONTEXT_ORDER_IDENTITY = "ae0500780a678072cc69176b9146aa0ad7978c71a5f257a3a29e7d4cd9147914" +ABLATION_CONTEXT_ORDER_IDENTITY = "193270b890cf80f35b59320a2780fbd2db2b13fa4cc21d46c3032845f5322cac" +ACCEPTED_BASELINE_ID = "accepted-p3-context-direct-cubeful-ridge-alpha10-full-v1" +ACCEPTED_BASELINE_MODEL_IDENTITY = "e1df7b51b8765de7b5457a64a6555c9a1e598ea2d4b4d06a897d6ee2df4543f5" +FULL_CONTEXT_MODEL_ID = "native-cubeful-p3-context-additive-ridge-v1" +FULL_CONTEXT_MODEL_IDENTITY = "ad426d22694783c674facd9722c713c44fe82c0de9151805e1155c8952382d69" +CANDIDATE_MODEL_ID = "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" +REMOVED_FIELD = "cubeful_crawford" +CONTEXT_FIELDS = tuple(item.feature_id for item in CUBEFUL_CONTEXT_REGISTRY) +ABLATION_CONTEXT_FIELDS = tuple(item for item in CONTEXT_FIELDS if item != REMOVED_FIELD) +SOURCE_FEATURE_IDS = tuple(item.feature_id for item in (*P3_REGISTRY, *CUBEFUL_CONTEXT_REGISTRY)) +RETAINED_FEATURE_IDS = tuple(item for item in SOURCE_FEATURE_IDS if item != REMOVED_FIELD) +SOURCE_FEATURE_INDEXES = tuple(index for index, item in enumerate(SOURCE_FEATURE_IDS) if item != REMOVED_FIELD) + + +def _sha(value: Any) -> str: + return hashlib.sha256(stable_json(value).encode()).hexdigest() + + +def _write(path: Path, value: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(stable_json(value, pretty=True), encoding="utf-8") + + +def _with_identity(payload: dict[str, Any]) -> dict[str, Any]: + payload["identity_sha256"] = _sha(payload) + return payload + + +def _identity_valid(payload: Mapping[str, Any]) -> bool: + value = dict(payload) + observed = value.pop("identity_sha256", None) + return observed == _sha(value) + + +def _utc_now() -> str: + return datetime.now(timezone.utc).isoformat().replace("+00:00", "Z") + + +def record_preflight(*, evidence_root: Path, phase: str, storage_path: Path) -> dict[str, Any]: + """Append a read-only shared-host capacity observation.""" + + memory: dict[str, int] = {} + for line in Path("/proc/meminfo").read_text().splitlines(): + key, raw = line.split(":", 1) + memory[key] = int(raw.strip().split()[0]) + stat = os.statvfs(storage_path) + payload = { + "recorded_at_utc": _utc_now(), + "phase": phase, + "host": socket.gethostname(), + "load_average": list(os.getloadavg()), + "logical_host_cpus": os.cpu_count(), + "process_affinity_cpus": len(os.sched_getaffinity(0)), + "mem_total_kib": memory["MemTotal"], + "mem_available_kib": memory["MemAvailable"], + "storage_path": str(storage_path.resolve()), + "storage_free_bytes": stat.f_bavail * stat.f_frsize, + "storage_free_inodes": stat.f_favail, + "nice": os.getpriority(os.PRIO_PROCESS, 0), + "numerical_thread_limit": 1, + "process_visibility": "sandbox PID namespace; host load/memory/filesystem are visible", + "protected_process_actions": [], + "disposition": "PASS_SUBSTANTIAL_HEADROOM", + } + evidence_root.mkdir(parents=True, exist_ok=True) + with (evidence_root / "preflight-log.jsonl").open("a", encoding="utf-8") as stream: + stream.write(stable_json(payload) + "\n") + return payload + + +def _verify_package(root: Path, expected_identity: str, generation: str) -> dict[str, str]: + manifest = json.loads((root / "manifest.json").read_text()) + if manifest["package_identity_sha256"] != expected_identity: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package identity") + files = {} + for item in manifest["files"]: + path = root / item["path"] + if not path.exists() or sha256_file(path) != item["sha256"]: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: {generation} package file") + files[item["path"]] = item["sha256"] + return files + + +def _full_context_model(generation_1_root: Path) -> dict[str, Any]: + models = json.loads((generation_1_root / "models.json").read_text()) + matches = [item for item in models["models"] if item["model_id"] == FULL_CONTEXT_MODEL_ID] + if len(matches) != 1 or matches[0]["model_identity_sha256"] != FULL_CONTEXT_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 full-context model") + return matches[0] + + +def freeze_receipts( + *, protocol_path: Path, split_manifest: Path, generation_1_root: Path, + generation_2_root: Path, generation_1_result: Path, generation_2_result: Path, + reference_models: Path, reference_development: Path, output_root: Path, +) -> dict[str, Any]: + """Verify and durably freeze every input authority before fitting/scoring.""" + + if sha256_file(protocol_path) != PROTOCOL_SHA256: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen Generation 3 protocol") + generation_1_files = _verify_package(generation_1_root, GENERATION_1_PACKAGE_IDENTITY, "Generation 1") + generation_2_files = _verify_package(generation_2_root, GENERATION_2_PACKAGE_IDENTITY, "Generation 2") + result_1 = json.loads(generation_1_result.read_text()) + result_2 = json.loads(generation_2_result.read_text()) + if result_1["identity_sha256"] != GENERATION_1_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 result") + if result_2["identity_sha256"] != GENERATION_2_RESULT_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 result") + if result_2["terminal_classification"] != "CONTEXT_FAILURE_LOCALIZED" or result_2["routing_cell"] != { + "family": "crawford_state", "cell": "money_or_zero", + }: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 2 routing") + + split = json.loads(split_manifest.read_text()) + authorities = json.loads((generation_1_root / "frozen-authorities.json").read_text()) + development = json.loads((generation_1_root / "development.json").read_text()) + cache = json.loads((generation_1_root / "training-cache.json").read_text()) + full_model = _full_context_model(generation_1_root) + source_authority = authorities["source_and_partition_authority"] + if sha256_file(split_manifest) != source_authority["split_manifest_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest hash") + if split["manifest_identity_sha256"] != source_authority["split_manifest_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: split manifest identity") + observed_partition = ( + source_authority["train"]["membership_sha256"], + source_authority["development"]["membership_sha256"], + source_authority["train"]["candidates"], + source_authority["development"]["candidates"], + source_authority["development"]["decisions"], + source_authority["development"]["complete_games"], + split["selection"]["holdout"]["game_membership_sha256"], + ) + expected_partition = ( + TRAIN_MEMBERSHIP_IDENTITY, DEVELOPMENT_MEMBERSHIP_IDENTITY, + EXPECTED_TRAIN_ROWS, EXPECTED_DEVELOPMENT_ROWS, + EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + DEVELOPMENT_MEMBERSHIP_IDENTITY, + ) + if observed_partition != expected_partition: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: TRAIN/DEVELOPMENT membership") + if development["fold_version"] != DEVELOPMENT_FOLD_VERSION or development["fold_seed"] != DEVELOPMENT_FOLD_SEED: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: grouped folds") + if authorities["target_authority"] != result_1["target"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: target mapping") + if authorities["feature_authority"]["position_registry_sha256"] != POSITION_REGISTRY_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: P3 registry") + if tuple(authorities["feature_authority"]["context_fields"]) != CONTEXT_FIELDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order") + if _sha(list(CONTEXT_FIELDS)) != FULL_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: full context order identity") + if _sha(list(ABLATION_CONTEXT_FIELDS)) != ABLATION_CONTEXT_ORDER_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: ablation context order identity") + baseline = authorities["model_comparison"]["baseline"] + if baseline["id"] != ACCEPTED_BASELINE_ID or baseline["model_identity_sha256"] != ACCEPTED_BASELINE_MODEL_IDENTITY: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted baseline") + if sha256_file(reference_models) != baseline["models_artifact_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline model artifact") + if sha256_file(reference_development) != baseline["accepted_development_evidence_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: baseline DEVELOPMENT artifact") + if cache["candidate_rows"] != EXPECTED_TRAIN_ROWS or cache["checkpoint"] != TRAIN_CHECKPOINT: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: accepted TRAIN cache") + + transform = full_model["transform"] + if tuple(transform["feature_ids"]) != SOURCE_FEATURE_IDS: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: Generation 1 additive feature order") + optimizer = full_model["optimizer"] + frozen_optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "alpha": 100.0, "batch_size": 32768, "beta1": 0.9, "beta2": 0.999, + "epochs": 12, "initial_learning_rate": 0.015, + "learning_rate_epoch_multiplier": 0.75, "seed": 20260823, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + } + for key, expected in frozen_optimizer.items(): + if optimizer[key] != expected: + raise RuntimeError(f"BLOCKED_INPUT_IDENTITY_MISMATCH: optimizer {key}") + + output_root.mkdir(parents=True, exist_ok=True) + context_receipt = _with_identity({ + "version": VERSION + "-context-order-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "full_context_fields": list(CONTEXT_FIELDS), + "full_context_order_sha256": FULL_CONTEXT_ORDER_IDENTITY, + "removed_field": REMOVED_FIELD, + "removed_full_context_index": CONTEXT_FIELDS.index(REMOVED_FIELD), + "retained_context_fields": list(ABLATION_CONTEXT_FIELDS), + "retained_context_count": len(ABLATION_CONTEXT_FIELDS), + "retained_context_order_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "substitute_crawford_encoding": None, + "interactions": 0, + }) + _write(output_root / "context-order-receipt.json", context_receipt) + model_receipt = _with_identity({ + "version": VERSION + "-model-authority-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "candidate_count": 1, + "candidate_model_id": CANDIDATE_MODEL_ID, + "control_model_id": FULL_CONTEXT_MODEL_ID, + "control_model_identity_sha256": FULL_CONTEXT_MODEL_IDENTITY, + "accepted_baseline_id": ACCEPTED_BASELINE_ID, + "accepted_baseline_model_identity_sha256": ACCEPTED_BASELINE_MODEL_IDENTITY, + "position_registry_identity_sha256": POSITION_REGISTRY_IDENTITY, + "source_feature_count": FULL_WIDTH, + "retained_feature_count": len(RETAINED_FEATURE_IDS), + "retained_feature_ids_sha256": _sha(list(RETAINED_FEATURE_IDS)), + "retained_source_feature_indexes": list(SOURCE_FEATURE_INDEXES), + "basis": authorities["model_comparison"]["additive_basis"], + "regularization_grid": [ALPHA], + "optimizer": frozen_optimizer, + "hyperparameter_searches": 0, + "feature_searches": 0, + "alternative_family_ablations": 0, + }) + _write(output_root / "model-authority-receipt.json", model_receipt) + protected_log = _with_identity({ + "version": VERSION + "-protected-access-log-v1", + "access_budget": 0, + "accesses": [], + "protected_membership_opened": False, + "protected_predictions_opened": False, + "protected_scored": False, + }) + _write(output_root / "protected-access-log.json", protected_log) + receipt = _with_identity({ + "version": VERSION + "-frozen-authorities-input-receipt-v1", + "status": "FROZEN_BEFORE_FIT_AND_DEVELOPMENT", + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "protocol": { + "path": str(protocol_path), "freeze_commit": PROTOCOL_FREEZE_COMMIT, + "sha256": PROTOCOL_SHA256, + }, + "generation_1": { + "result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "package_files_sha256": generation_1_files, + }, + "generation_2": { + "result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "package_files_sha256": generation_2_files, + "terminal_classification": result_2["terminal_classification"], + "routing_cell": result_2["routing_cell"], + }, + "source": { + "authority": source_authority["source"], + "source_manifest_sha256": source_authority["source_manifest_sha256"], + "split_manifest_path": str(split_manifest), + "split_manifest_sha256": sha256_file(split_manifest), + "split_manifest_identity_sha256": split["manifest_identity_sha256"], + "training_cache_receipt_sha256": sha256_file(generation_1_root / "training-cache.json"), + "training_cache_deterministic_identity_sha256": cache["deterministic_identity_sha256"], + }, + "partitions": { + "train": source_authority["train"], + "development": source_authority["development"], + "protected_final_evaluation": { + "access_budget": 0, "membership_opened": False, + "predictions_opened": False, "scored": False, + "path_recorded": False, + }, + }, + "target": authorities["target_authority"], + "grouped_folds": authorities["development_grouped_folds"], + "context_order_receipt": { + "path": "context-order-receipt.json", + "sha256": sha256_file(output_root / "context-order-receipt.json"), + "identity_sha256": context_receipt["identity_sha256"], + }, + "model_authority_receipt": { + "path": "model-authority-receipt.json", + "sha256": sha256_file(output_root / "model-authority-receipt.json"), + "identity_sha256": model_receipt["identity_sha256"], + }, + "activity_boundary": authorities["activity_boundary"], + "experiment_limits": { + "model_fits": 1, "candidate_count": 1, "hyperparameter_searches": 0, + "feature_searches": 0, "alternative_family_ablations": 0, + "protected_access_budget": 0, + }, + }) + _write(output_root / "frozen-authorities-input-receipt.json", receipt) + return receipt + + +@dataclass +class AblationTransform: + feature_ids: tuple[str, ...] + source_feature_indexes: tuple[int, ...] + mean: np.ndarray + scale: np.ndarray + knots: np.ndarray + + @property + def width(self) -> int: + return len(self.feature_ids) + + @property + def basis_width(self) -> int: + return self.width * 4 + + def basis_float32(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=np.float32) + z = (selected - self.mean.astype(np.float32)) / self.scale.astype(np.float32) + output = np.empty((len(z), self.width, 4), dtype=np.float32) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots.astype(np.float32)[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def basis(self, source_x: np.ndarray) -> np.ndarray: + selected = np.asarray(source_x[:, self.source_feature_indexes], dtype=float) + z = (selected - self.mean) / self.scale + output = np.empty((len(z), self.width, 4), dtype=float) + output[:, :, 0] = z + output[:, :, 1:] = np.maximum(z[:, :, None] - self.knots[None, :, :], 0.0) + return output.reshape(len(z), -1) + + def descriptor(self) -> dict[str, Any]: + return { + "feature_ids": list(self.feature_ids), + "source_feature_indexes": list(self.source_feature_indexes), + "standard_scaler_mean": self.mean.tolist(), + "standard_scaler_scale": self.scale.tolist(), + "hinge_quantiles": [0.25, 0.5, 0.75], + "hinge_knots_standardized": self.knots.tolist(), + "basis_order": "feature-major standardized linear,q25,q50,q75 positive hinges", + "interactions": 0, + "removed_before_basis_construction": REMOVED_FIELD, + } + + +@dataclass +class AblationModel: + transform: AblationTransform + coefficients: np.ndarray + intercept: float + optimizer: dict[str, Any] + + def predict(self, source_x: np.ndarray) -> np.ndarray: + return self.transform.basis(source_x) @ self.coefficients + self.intercept + + def grouped_contributions(self, source_x: np.ndarray) -> np.ndarray: + basis = self.transform.basis(source_x) + return (basis * self.coefficients).reshape(len(source_x), self.transform.width, 4).sum(axis=2) + + def descriptor(self) -> dict[str, Any]: + payload = { + "model_id": CANDIDATE_MODEL_ID, + "target": TARGET, + "alpha": ALPHA, + "training_checkpoint": TRAIN_CHECKPOINT, + "transform": self.transform.descriptor(), + "coefficients": self.coefficients.tolist(), + "intercept": self.intercept, + "optimizer": self.optimizer, + "exact_explanation_scale": "GNU native Cubeful equity in normalized static next-player-on-roll perspective", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + } + payload["model_identity_sha256"] = _sha(payload) + return payload + + +def _candidate_transform(generation_1_root: Path) -> AblationTransform: + full = _full_context_model(generation_1_root)["transform"] + keep = np.asarray(SOURCE_FEATURE_INDEXES, dtype=int) + return AblationTransform( + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + np.asarray(full["standard_scaler_mean"], dtype=float)[keep], + np.asarray(full["standard_scaler_scale"], dtype=float)[keep], + np.asarray(full["hinge_knots_standardized"], dtype=float)[keep], + ) + + +def fit_candidate(*, cache_root: Path, generation_1_root: Path, receipt_path: Path, + output_path: Path) -> dict[str, Any]: + """Fit the one permitted candidate without opening DEVELOPMENT.""" + + receipt = json.loads(receipt_path.read_text()) + if not _identity_valid(receipt) or receipt["status"] != "FROZEN_BEFORE_FIT_AND_DEVELOPMENT": + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: frozen receipt") + parts = load_cache(cache_root) + transform = _candidate_transform(generation_1_root) + rows = sum(len(part.indexes) for part in parts) + target_mean = sum(float(np.asarray(part.y[part.indexes], dtype=float).sum()) for part in parts) / rows + parameter = np.zeros(transform.basis_width + 1, dtype=np.float32) + parameter[-1] = target_mean + first = np.zeros_like(parameter) + second = np.zeros_like(parameter) + step = 0 + epoch_records = [] + started = time.time() + for epoch in range(EPOCHS): + sum_squared = 0.0 + seen = 0 + maximum_update = 0.0 + for source_x, y in _iter_cache(parts): + basis = transform.basis_float32(source_x) + residual = basis @ parameter[:-1] + parameter[-1] - y + sum_squared += float(np.square(residual).sum()) + gradient = np.append(residual @ basis / len(source_x), residual.mean()).astype(np.float32) + gradient[:-1] += (ALPHA / rows) * parameter[:-1] + learning_rate = 0.015 * (0.75 ** epoch) * min(1.0, len(source_x) / BATCH_SIZE) + maximum_update = max( + maximum_update, + _adam_step(parameter, gradient, first, second, step + 1, learning_rate), + ) + step += 1 + seen += len(source_x) + epoch_records.append({ + "epoch": epoch + 1, + "online_rmse": math.sqrt(sum_squared / seen), + "maximum_absolute_parameter_update": float(maximum_update), + }) + optimizer = { + "algorithm": "deterministic streaming Adam on Ridge objective", + "objective": "0.5*sum_squared_error + 0.5*alpha*squared_coefficient_norm", + "alpha": ALPHA, "epochs": EPOCHS, "batch_size": BATCH_SIZE, + "initial_learning_rate": 0.015, "learning_rate_epoch_multiplier": 0.75, + "beta1": 0.9, "beta2": 0.999, "seed": 20260823, + "training_rows": rows, "updates": step, "epoch_records": epoch_records, + "training_arithmetic": "float32 basis/optimizer; float64 retained inference/reconstruction", + "candidate_index": 0, + } + model = AblationModel(transform, parameter[:-1].astype(float), float(parameter[-1]), optimizer) + payload = _with_identity({ + "version": VERSION + "-candidate-model-v1", + "status": "PASS_SINGLE_CANDIDATE_FIT_BEFORE_DEVELOPMENT", + "candidate_count": 1, + "model": model.descriptor(), + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "train_membership_identity_sha256": TRAIN_MEMBERSHIP_IDENTITY, + "development_accesses": 0, + "protected_accesses": 0, + "hyperparameter_searches": 0, + "feature_searches": 0, + }) + _write(output_path, payload) + return payload + + +def load_candidate(path: Path) -> AblationModel: + payload = json.loads(path.read_text()) + item = payload["model"] + transform = item["transform"] + model = AblationModel( + AblationTransform( + tuple(transform["feature_ids"]), tuple(transform["source_feature_indexes"]), + np.asarray(transform["standard_scaler_mean"]), + np.asarray(transform["standard_scaler_scale"]), + np.asarray(transform["hinge_knots_standardized"]), + ), + np.asarray(item["coefficients"]), float(item["intercept"]), dict(item["optimizer"]), + ) + if model.descriptor()["model_identity_sha256"] != item["model_identity_sha256"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: candidate model") + return model + + +class DevelopmentMetrics: + def __init__(self) -> None: + self.global_metric = RegressionMetrics() + self.folds = {index: RegressionMetrics() for index in range(4)} + + def add(self, prediction: np.ndarray, truth: np.ndarray, folds: np.ndarray) -> None: + self.global_metric.add(prediction, truth) + for fold in range(4): + selected = folds == fold + if np.any(selected): + self.folds[fold].add(prediction[selected], truth[selected]) + + def result(self) -> dict[str, Any]: + return { + "global": self.global_metric.result(), + "development_complete_game_folds": { + str(index): self.folds[index].result() for index in range(4) + }, + } + + +def _comparison(candidate: Mapping[str, Any], control: Mapping[str, Any]) -> dict[str, Any]: + cand_global, control_global = candidate["global"], control["global"] + lower_rmse = lower_mae = lower_both = 0 + max_rmse_regression = -math.inf + max_mae_regression = -math.inf + folds = {} + for fold in range(4): + key = str(fold) + cand_fold = candidate["development_complete_game_folds"][key] + control_fold = control["development_complete_game_folds"][key] + rmse_recovery = control_fold["rmse"] - cand_fold["rmse"] + mae_recovery = control_fold["mae"] - cand_fold["mae"] + rmse_lower = rmse_recovery > 0 + mae_lower = mae_recovery > 0 + lower_rmse += int(rmse_lower) + lower_mae += int(mae_lower) + lower_both += int(rmse_lower and mae_lower) + max_rmse_regression = max(max_rmse_regression, -rmse_recovery) + max_mae_regression = max(max_mae_regression, -mae_recovery) + folds[key] = { + "rmse_recovery": rmse_recovery, "mae_recovery": mae_recovery, + "candidate_lower_rmse": rmse_lower, "candidate_lower_mae": mae_lower, + } + return { + "global": { + "rmse_recovery": control_global["rmse"] - cand_global["rmse"], + "mae_recovery": control_global["mae"] - cand_global["mae"], + "absolute_bias_change": abs(cand_global["bias"]) - abs(control_global["bias"]), + "signed_bias_change": cand_global["bias"] - control_global["bias"], + "correlation_change": cand_global["correlation"] - control_global["correlation"], + "r2_change": cand_global["r2"] - control_global["r2"], + }, + "fold_counts": { + "lower_rmse": lower_rmse, "lower_mae": lower_mae, "lower_both": lower_both, + }, + "maximum_fold_rmse_regression": max_rmse_regression, + "maximum_fold_mae_regression": max_mae_regression, + "folds": folds, + } + + +def evaluate_decision(control_comparison: Mapping[str, Any], + baseline_comparison: Mapping[str, Any]) -> dict[str, Any]: + """Apply the prospectively frozen causal and secondary rules.""" + + recovery = control_comparison["global"] + causal_checks = { + "global_rmse_recovery_at_least_0_005": recovery["rmse_recovery"] >= 0.005, + "global_mae_recovery_at_least_0_002": recovery["mae_recovery"] >= 0.002, + "lower_both_in_at_least_3_of_4_folds": control_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": control_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": control_comparison["maximum_fold_mae_regression"] <= 0.002, + "absolute_global_bias_worsening_at_most_0_005": recovery["absolute_bias_change"] <= 0.005, + } + baseline = baseline_comparison["global"] + secondary_checks = { + "global_rmse_improvement_at_least_0_002": baseline["rmse_recovery"] >= 0.002, + "global_mae_improvement_at_least_0_002": baseline["mae_recovery"] >= 0.002, + "absolute_global_bias_worsening_at_most_0_005": baseline["absolute_bias_change"] <= 0.005, + "lower_both_in_at_least_3_of_4_folds": baseline_comparison["fold_counts"]["lower_both"] >= 3, + "maximum_fold_rmse_regression_at_most_0_002": baseline_comparison["maximum_fold_rmse_regression"] <= 0.002, + "maximum_fold_mae_regression_at_most_0_002": baseline_comparison["maximum_fold_mae_regression"] <= 0.002, + } + causal = all(causal_checks.values()) + beats_baseline = all(secondary_checks.values()) + return { + "classification": "CRAWFORD_CONTEXT_FAMILY_CAUSAL" if causal else "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL", + "causal_rule_checks": causal_checks, + "secondary_label": "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" if beats_baseline else None, + "secondary_rule_checks": secondary_checks, + "production": "UNCHANGED", + "candidate_disposition": "CANDIDATE_FOR_FUTURE_INDEPENDENT_MODEL_SELECTION" if beats_baseline else None, + } + + +def score_development( + *, shallow_root: Path, split_manifest: Path, generation_1_root: Path, + receipt_path: Path, model_path: Path, output_root: Path, +) -> dict[str, Any]: + """Score the one durable candidate on DEVELOPMENT and apply frozen rules.""" + + if not model_path.exists(): + raise RuntimeError("candidate model must be durable before DEVELOPMENT access") + receipt = json.loads(receipt_path.read_text()) + protected = json.loads((output_root / "protected-access-log.json").read_text()) + if not _identity_valid(receipt) or not _identity_valid(protected) or protected["accesses"]: + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: receipt/protected log") + model = load_candidate(model_path) + split = json.loads(split_manifest.read_text()) + _, holdout, _ = _membership(split) + campaign = str(split["source_authority"]["campaign"]) + metrics = DevelopmentMetrics() + decisions: set[str] = set() + groups: set[str] = set() + samples: tuple[np.ndarray, list[str], np.ndarray] | None = None + rows = 0 + started = time.time() + for path in _candidate_files(shallow_root): + games = holdout.get(_partition_key(path), set()) + if not games: + continue + host, worker = _partition_key(path) + for batch in pq.ParquetFile(path).iter_batches(batch_size=8192, columns=list(SOURCE_COLUMNS)): + data = batch.to_pydict() + chosen = np.flatnonzero(np.fromiter((str(value) in games for value in data["game_key"]), dtype=bool)) + if not len(chosen): + continue + positions = [str(data["static_position_id_on_roll"][index]) for index in chosen] + match_ids = [str(data["source_match_id"][index]) for index in chosen] + source_x = np.column_stack((position_feature_matrix(positions), cubeful_context_matrix(match_ids))) + truth = -np.asarray(data["native_equity"], dtype=float)[chosen] + group_ids = ["\0".join((campaign, host, worker, str(data["game_key"][index]))) for index in chosen] + folds = np.fromiter((_fold(group_id) for group_id in group_ids), dtype=np.int8) + prediction = model.predict(source_x) + metrics.add(prediction, truth, folds) + if samples is None: + count = min(2, len(chosen)) + samples = ( + source_x[:count].copy(), + [str(data["decision_id"][index]) for index in chosen[:count]], + truth[:count].copy(), + ) + rows += len(chosen) + decisions.update(str(data["decision_id"][index]) for index in chosen) + groups.update(group_ids) + if (rows, len(decisions), len(groups)) != ( + EXPECTED_DEVELOPMENT_ROWS, EXPECTED_DEVELOPMENT_DECISIONS, DEVELOPMENT_COMPLETE_GAMES, + ): + raise RuntimeError("BLOCKED_INPUT_IDENTITY_MISMATCH: DEVELOPMENT population") + if samples is None: + raise RuntimeError("DEVELOPMENT sample missing") + + generation_1_development = json.loads((generation_1_root / "development.json").read_text()) + baseline = generation_1_development["metrics"]["accepted_baseline"] + control = generation_1_development["metrics"][FULL_CONTEXT_MODEL_ID] + candidate = metrics.result() + controls = { + "accepted_baseline": baseline, + FULL_CONTEXT_MODEL_ID: control, + CANDIDATE_MODEL_ID: candidate, + } + development_payload = _with_identity({ + "version": VERSION + "-development-metrics-v1", + "status": "PASS_SINGLE_DEVELOPMENT_SCORE", + "access": "DEVELOPMENT/ADAPTIVE_SELECTION", + "candidates": rows, "decisions": len(decisions), "complete_game_groups": len(groups), + "membership_identity_sha256": DEVELOPMENT_MEMBERSHIP_IDENTITY, + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "metrics": controls, + "control_source": "exact frozen Generation 1 DEVELOPMENT metrics; controls not refit", + "protected_accesses": 0, + }) + _write(output_root / "development-metrics.json", development_payload) + fold_payload = _with_identity({ + "version": VERSION + "-fold-metrics-v1", + "status": "PASS", + "fold_version": DEVELOPMENT_FOLD_VERSION, "fold_seed": DEVELOPMENT_FOLD_SEED, + "models": { + name: value["development_complete_game_folds"] for name, value in controls.items() + }, + }) + _write(output_root / "fold-metrics.json", fold_payload) + control_comparison = _comparison(candidate, control) + baseline_comparison = _comparison(candidate, baseline) + comparison_payload = _with_identity({ + "version": VERSION + "-recovery-comparison-v1", + "status": "PASS", + "candidate_model_id": CANDIDATE_MODEL_ID, + "versus_generation_1_full_context_control": control_comparison, + "versus_accepted_baseline": baseline_comparison, + }) + _write(output_root / "recovery-comparison.json", comparison_payload) + decision = evaluate_decision(control_comparison, baseline_comparison) + terminal_payload = _with_identity({ + "version": VERSION + "-terminal-classification-v1", + "status": "PASS_DETERMINISTIC_CLASSIFICATION", + **decision, + "generation_2_localization_disposition": ( + "CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" if decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + else "NON_CAUSAL_UNDER_SINGLE_FAMILY_FOLLOWUP" + ), + "protected_access_authorized": False, + "next_experiment": None, + }) + _write(output_root / "terminal-classification.json", terminal_payload) + + sample_x, decision_ids, sample_truth = samples + prediction = model.predict(sample_x) + contributions = model.grouped_contributions(sample_x) + reconstructed = contributions.sum(axis=1) + model.intercept + maximum_error = float(np.max(np.abs(prediction - reconstructed))) + if maximum_error > 1e-10: + raise RuntimeError("candidate additive reconstruction failed") + explanation_payload = _with_identity({ + "version": VERSION + "-contribution-evidence-v1", "status": "PASS", + "definition": "prediction = intercept + sum of grouped per-feature linear/hinge contributions", + "model_id": CANDIDATE_MODEL_ID, + "model_identity_sha256": model.descriptor()["model_identity_sha256"], + "decision_ids": decision_ids, "truth": sample_truth.tolist(), + "prediction": prediction.tolist(), "intercept": model.intercept, + "feature_contributions": [ + {"feature_id": feature_id, "row_values": contributions[:, index].tolist()} + for index, feature_id in enumerate(model.transform.feature_ids) + ], + "global_maximum_absolute_reconstruction_error": maximum_error, + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + }) + _write(output_root / "contribution-evidence.json", explanation_payload) + run_record = _with_identity({ + "version": VERSION + "-run-record-v1", + "status": "PASS_COMPLETE_SINGLE_CANDIDATE_DEVELOPMENT_EXPERIMENT", + "frozen_receipt_sha256": sha256_file(receipt_path), + "candidate_model_sha256": sha256_file(model_path), + "candidate_model_identity_sha256": model.descriptor()["model_identity_sha256"], + "development_metrics_identity_sha256": development_payload["identity_sha256"], + "terminal_classification_identity_sha256": terminal_payload["identity_sha256"], + "elapsed_seconds": time.time() - started, + "peak_rss_kib": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss, + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "alternative_family_ablations": 0, "protected_accesses": 0, + "new_gnu": 0, "new_sage": 0, "new_matches": 0, "new_labels": 0, + "generic_new_0ply": 0, "sage_gnu_campaign_training_rows": 0, + }) + _write(output_root / "run-record.json", run_record) + return terminal_payload + + +def build_summary(*, evidence_root: Path, result_path: Path) -> dict[str, Any]: + receipt = json.loads((evidence_root / "frozen-authorities-input-receipt.json").read_text()) + model = json.loads((evidence_root / "candidate-model.json").read_text()) + development = json.loads((evidence_root / "development-metrics.json").read_text()) + comparison = json.loads((evidence_root / "recovery-comparison.json").read_text()) + terminal = json.loads((evidence_root / "terminal-classification.json").read_text()) + protected = json.loads((evidence_root / "protected-access-log.json").read_text()) + payload = _with_identity({ + "version": VERSION + "-result-v1", + "status": "PASS_COMPLETE", + "terminal_classification": terminal["classification"], + "secondary_label": terminal["secondary_label"], + "generation_2_localization_disposition": terminal["generation_2_localization_disposition"], + "starting_implementation_head": STARTING_IMPLEMENTATION_HEAD, + "generation_1_result_identity_sha256": GENERATION_1_RESULT_IDENTITY, + "generation_1_package_identity_sha256": GENERATION_1_PACKAGE_IDENTITY, + "generation_2_result_identity_sha256": GENERATION_2_RESULT_IDENTITY, + "generation_2_package_identity_sha256": GENERATION_2_PACKAGE_IDENTITY, + "train_membership_identity_sha256": receipt["partitions"]["train"]["membership_sha256"], + "development_membership_identity_sha256": receipt["partitions"]["development"]["membership_sha256"], + "target": receipt["target"], + "candidate_model_id": CANDIDATE_MODEL_ID, + "candidate_model_identity_sha256": model["model"]["model_identity_sha256"], + "context_order_identity_sha256": ABLATION_CONTEXT_ORDER_IDENTITY, + "development_global_metrics": { + name: value["global"] for name, value in development["metrics"].items() + }, + "recovery": comparison["versus_generation_1_full_context_control"], + "accepted_baseline_comparison": comparison["versus_accepted_baseline"], + "causal_rule_checks": terminal["causal_rule_checks"], + "secondary_rule_checks": terminal["secondary_rule_checks"], + "model_fits": 1, "candidate_count": 1, + "hyperparameter_searches": 0, "feature_searches": 0, + "protected_access_count": len(protected["accesses"]), + "accepted_product_architecture": "ridge-ranking-hadd-value-explanation-sidecar-v1", + "pairwise_ridge_role": "SOLE_RECOMMENDATION_RANKING_AUTHORITY_UNCHANGED", + "compact_hadd_role": "ACCEPTED_PROBABILITY_VALUE_EXPLANATION_AUTHORITY_UNCHANGED", + "calculated_cubeful": "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED", + "production": "UNCHANGED", "analyzer": "UNCHANGED", + "canonical": "UNCHANGED", "corpus": "UNCHANGED", + "candidate_disposition": terminal["candidate_disposition"], + "next_task_status": "WAITING_FOR_RESEARCH_DIRECTOR", + "next_experiment": None, + }) + _write(evidence_root / "result-summary.json", payload) + _write(result_path, payload) + return payload + + +def verify_determinism(evidence_root: Path) -> dict[str, Any]: + names = ( + "context-order-receipt.json", "model-authority-receipt.json", + "protected-access-log.json", "frozen-authorities-input-receipt.json", + "candidate-model.json", "development-metrics.json", "fold-metrics.json", + "recovery-comparison.json", "terminal-classification.json", + "contribution-evidence.json", "run-record.json", "result-summary.json", + "test-evidence.json", + ) + checks = [] + payloads = {} + for name in names: + path = evidence_root / name + value = json.loads(path.read_text()) if path.exists() else {} + payloads[name] = value + checks.append({ + "name": "internal_identity:" + name, + "status": "PASS" if value and _identity_valid(value) else "FAIL", + }) + comparison = payloads["recovery-comparison.json"] + rerun = evaluate_decision( + comparison["versus_generation_1_full_context_control"], + comparison["versus_accepted_baseline"], + ) + terminal = payloads["terminal-classification.json"] + checks.extend(( + { + "name": "deterministic_terminal_classification", + "status": "PASS" if all(terminal[key] == value for key, value in rerun.items()) else "FAIL", + }, + { + "name": "exactly_one_model_fit", + "status": "PASS" if payloads["run-record.json"]["model_fits"] == 1 and payloads["run-record.json"]["candidate_count"] == 1 else "FAIL", + }, + { + "name": "crawford_only_removed_before_basis", + "status": "PASS" if payloads["candidate-model.json"]["model"]["transform"]["removed_before_basis_construction"] == REMOVED_FIELD and REMOVED_FIELD not in payloads["candidate-model.json"]["model"]["transform"]["feature_ids"] else "FAIL", + }, + { + "name": "protected_access_exactly_zero", + "status": "PASS" if payloads["protected-access-log.json"]["access_budget"] == 0 and not payloads["protected-access-log.json"]["accesses"] else "FAIL", + }, + { + "name": "accepted_architecture_unchanged", + "status": "PASS" if payloads["result-summary.json"]["accepted_product_architecture"] == "ridge-ranking-hadd-value-explanation-sidecar-v1" else "FAIL", + }, + { + "name": "production_analyzer_canonical_corpus_unchanged", + "status": "PASS" if all(payloads["result-summary.json"][key] == "UNCHANGED" for key in ("production", "analyzer", "canonical", "corpus")) else "FAIL", + }, + )) + payload = _with_identity({ + "version": VERSION + "-deterministic-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "checks": checks, + "verified_artifact_sha256": {name: sha256_file(evidence_root / name) for name in names}, + "python": platform.python_version(), + }) + _write(evidence_root / "deterministic-verification.json", payload) + return payload + + +def build_manifest(evidence_root: Path) -> dict[str, Any]: + excluded = {"manifest.json", "SHA256SUMS", "self-verification.json"} + files = sorted(path for path in evidence_root.iterdir() if path.is_file() and path.name not in excluded) + entries = [ + {"path": path.name, "sha256": sha256_file(path), "size_bytes": path.stat().st_size} + for path in files + ] + payload = _with_identity({"version": VERSION + "-manifest-v1", "files": entries}) + payload["package_identity_sha256"] = payload.pop("identity_sha256") + _write(evidence_root / "manifest.json", payload) + (evidence_root / "SHA256SUMS").write_text( + "".join(f"{item['sha256']} {item['path']}\n" for item in entries), encoding="utf-8", + ) + return payload + + +def verify_artifact_package(evidence_root: Path) -> dict[str, Any]: + manifest = json.loads((evidence_root / "manifest.json").read_text()) + value = dict(manifest) + package_identity = value.pop("package_identity_sha256") + checks = [{ + "name": "manifest_identity", + "status": "PASS" if package_identity == _sha(value) else "FAIL", + }] + for item in manifest["files"]: + path = evidence_root / item["path"] + checks.append({ + "name": "hash:" + item["path"], + "status": "PASS" if path.exists() and sha256_file(path) == item["sha256"] else "FAIL", + }) + deterministic = json.loads((evidence_root / "deterministic-verification.json").read_text()) + checks.append({"name": "deterministic_verification", "status": deterministic["status"]}) + payload = { + "version": VERSION + "-self-verification-v1", + "status": "PASS" if all(item["status"] == "PASS" for item in checks) else "FAIL", + "package_identity_sha256": package_identity, + "checks": checks, + "python": platform.python_version(), + } + _write(evidence_root / "self-verification.json", payload) + return payload diff --git a/tests/test_native_cubeful_crawford_ablation.py b/tests/test_native_cubeful_crawford_ablation.py new file mode 100644 index 0000000000000000000000000000000000000000..8f6bbbd7e90e97c706c8e5f9a0909a6cfc447401 --- /dev/null +++ b/tests/test_native_cubeful_crawford_ablation.py @@ -0,0 +1,74 @@ +from __future__ import annotations + +from pathlib import Path + +import numpy as np + +from backgammon_explainer.native_cubeful_crawford_ablation import ( + ABLATION_CONTEXT_FIELDS, + ABLATION_CONTEXT_ORDER_IDENTITY, + CANDIDATE_MODEL_ID, + CONTEXT_FIELDS, + REMOVED_FIELD, + RETAINED_FEATURE_IDS, + SOURCE_FEATURE_INDEXES, + AblationTransform, + _sha, + evaluate_decision, +) + + +def _comparison(*, rmse: float, mae: float, bias: float, lower_both: int, + max_rmse: float, max_mae: float) -> dict: + return { + "global": { + "rmse_recovery": rmse, "mae_recovery": mae, + "absolute_bias_change": bias, + }, + "fold_counts": {"lower_both": lower_both}, + "maximum_fold_rmse_regression": max_rmse, + "maximum_fold_mae_regression": max_mae, + } + + +def test_exact_single_crawford_field_ablation_and_order_identity() -> None: + assert len(CONTEXT_FIELDS) == 15 + assert len(ABLATION_CONTEXT_FIELDS) == 14 + assert REMOVED_FIELD == "cubeful_crawford" + assert CONTEXT_FIELDS.index(REMOVED_FIELD) == 12 + assert REMOVED_FIELD not in ABLATION_CONTEXT_FIELDS + assert REMOVED_FIELD not in RETAINED_FEATURE_IDS + assert len(SOURCE_FEATURE_INDEXES) == len(RETAINED_FEATURE_IDS) == 365 + assert _sha(list(ABLATION_CONTEXT_FIELDS)) == ABLATION_CONTEXT_ORDER_IDENTITY + assert CANDIDATE_MODEL_ID == "native-cubeful-p3-context-minus-crawford-additive-ridge-v1" + + +def test_ablation_transform_selects_before_basis_and_reconstructs() -> None: + transform = AblationTransform( + ("a", "c"), (0, 2), np.asarray([0.0, 0.0]), np.asarray([1.0, 1.0]), + np.asarray([[0.0, 0.5, 1.0], [0.0, 0.5, 1.0]]), + ) + source = np.asarray([[2.0, 999.0, 3.0]]) + basis = transform.basis(source) + assert basis.shape == (1, 8) + assert basis.tolist()[0] == [2.0, 2.0, 1.5, 1.0, 3.0, 3.0, 2.5, 2.0] + + +def test_prospective_causal_rule_requires_every_gate() -> None: + passing = _comparison(rmse=0.005, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + secondary = _comparison(rmse=0.002, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + decision = evaluate_decision(passing, secondary) + assert decision["classification"] == "CRAWFORD_CONTEXT_FAMILY_CAUSAL" + assert decision["secondary_label"] == "ABLATED_MODEL_BEATS_ACCEPTED_BASELINE" + failing = _comparison(rmse=0.004999, mae=0.002, bias=0.005, lower_both=3, max_rmse=0.002, max_mae=0.002) + assert evaluate_decision(failing, secondary)["classification"] == "CRAWFORD_CONTEXT_FAMILY_NOT_CAUSAL" + + +def test_no_protected_reader_or_product_mutation_phase() -> None: + module = Path("src/backgammon_explainer/native_cubeful_crawford_ablation.py").read_text() + runner = Path("scripts/run_native_cubeful_crawford_ablation.py").read_text() + assert "load_frozen_deep_rows" not in module + assert "score_protected" not in module + assert '"score-protected"' not in runner + assert "CUBEFUL_CALCULATION_AUTHORITY_BLOCKED" in module + assert '"production": "UNCHANGED"' in module ERROR: You've hit your usage limit. Visit https://chatgpt.com/codex/settings/usage to purchase more credits or try again at Aug 30th, 2026 8:41 PM. ERROR: You've hit your usage limit. Visit https://chatgpt.com/codex/settings/usage to purchase more credits or try again at Aug 30th, 2026 8:41 PM. tokens used 1,408,000